Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “protein function annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Protein classification using ontology classification.

MOTIVATION: The classification of proteins expressed by an organism is an important step in understanding the molecular biology of that organism. Traditionally, this classification has been performed by human experts. Human knowledge can recognise the functional properties that are sufficient to place an individual gene product into a particular protein family group. Automation of this task usually fails to meet the 'gold standard' of the human annotator because of the difficult recognition stage. The growing number of genomes, the rapid changes in knowledge and the central role of classification in the annotation process, however, motivates the need to automate this process. RESULTS: We capture human understanding of how to recognise members of the protein phosphatases family by domain architecture as an ontology. By describing protein instances in terms of the domains they contain, it is possible to use description logic reasoners and our ontology to assign those proteins to a protein family class. We have tested our system on classifying the protein phosphatases of the human and Aspergillus fumigatus genomes and found that our knowledge-based, automatic classification matches, and sometimes surpasses, that of the human annotators. We have made the classification process fast and reproducible and, where appropriate knowledge is available, the method can potentially be generalised for use with any protein family. AVAILABILITY: All components described in this paper are freely available. OWL ontology http://www.bioinf.man.ac.uk/phosphabase myGrid http://www.mygrid.org.uk Instance Store http://instancestore.man.ac.uk.

Algorithms↗

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Arabidopsis thaliana is the most widely-studied plant today. The concerted efforts of over 11 000 researchers and 4000 organizations around the world are generating a rich diversity and quantity of information and materials. This information is made available through a comprehensive on-line resource called the Arabidopsis Information Resource (TAIR) (http://arabidopsis.org), which is accessible via commonly used web browsers and can be searched and downloaded in a number of ways. In the last two years, efforts have been focused on increasing data content and diversity, functionally annotating genes and gene products with controlled vocabularies, and improving data retrieval, analysis and visualization tools. New information include sequence polymorphisms including alleles, germplasms and phenotypes, Gene Ontology annotations, gene families, protein information, metabolic pathways, gene expression data from microarray experiments and seed and DNA stocks. New data visualization and analysis tools include SeqViewer, which interactively displays the genome from the whole chromosome down to 10 kb of nucleotide sequence and AraCyc, a metabolic pathway database and map tool that allows overlaying expression data onto the pathway diagrams. Finally, we have recently incorporated seed and DNA stock information from the Arabidopsis Biological Resource Center (ABRC) and implemented a shopping-cart style on-line ordering system.

Arabidopsis↗

The angiotensin-converting enzyme (ACE) gene family of Anopheles gambiae.

BACKGROUND: Members of the M2 family of peptidases, related to mammalian angiotensin converting enzyme (ACE), play important roles in regulating a number of physiological processes. As more invertebrate genomes are sequenced, there is increasing evidence of a variety of M2 peptidase genes, even within a single species. The function of these ACE-like proteins is largely unknown. Sequencing of the A. gambiae genome has revealed a number of ACE-like genes but probable errors in the Ensembl annotation have left the number of ACE-like genes, and their structure, unclear. RESULTS: TBLASTN and sequence analysis of cDNAs revealed that the A. gambiae genome contains nine genes (AnoACE genes) which code for proteins with similarity to mammalian ACE. Eight of these genes code for putative single domain enzymes similar to other insect ACEs described so far. AnoACE9, however, has several features in common with mammalian somatic ACE such as a two domain structure and a hydrophobic C terminus. Four of the AnoACE genes (2, 3, 7 and 9) were shown to be expressed at a variety of developmental stages. Expression of AnoACE3, AnoACE7 and AnoACE9 is induced by a blood meal, with AnoACE7 showing the largest (approximately 10-fold) induction. CONCLUSION: Genes coding for two-domain ACEs have arisen several times during the course of evolution suggesting a common selective advantage to having an ACE with two active-sites in tandem in a single protein. AnoACE7 belongs to a sub-group of insect ACEs which are likely to be membrane-bound and which have an unusual, conserved gene structure.

Amino Acid Sequence↗

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF), Martinsried, near Munich, Germany, continues its longstanding tradition to develop and maintain high quality curated genome databases. In addition, efforts have been intensified to cover the wealth of complete genome sequences in a systematic, comprehensive form. Bioinformatics, supporting national as well as European sequencing and functional analysis projects, has resulted in several up-to-date genome-oriented databases. This report describes growing databases reflecting the progress of sequencing the Arabidopsis thaliana (MATDB) and Neurospora crassa genomes (MNCDB), the yeast genome database (MYGD) extended by functional analysis data, the database of annotated human EST-clusters (HIB) and the database of the complete cDNA sequences from the DHGP (German Human Genome Project). It also contains information on the up-to-date database of complete genomes (PEDANT), the classification of protein sequences (ProtFam) and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database. These databases can be accessed through the MIPS WWW server (http://www. mips.biochem.mpg.de).

Arabidopsis↗

The transcriptome of the sea urchin embryo.

The sea urchin Strongylocentrotus purpuratus is a model organism for study of the genomic control circuitry underlying embryonic development. We examined the complete repertoire of genes expressed in the S. purpuratus embryo, up to late gastrula stage, by means of high-resolution custom tiling arrays covering the whole genome. We detected complete spliced structures even for genes known to be expressed at low levels in only a few cells. At least 11,000 to 12,000 genes are used in embryogenesis. These include most of the genes encoding transcription factors and signaling proteins, as well as some classes of general cytoskeletal and metabolic proteins, but only a minor fraction of genes encoding immune functions and sensory receptors. Thousands of small asymmetric transcripts of unknown function were also detected in intergenic regions throughout the genome. The tiling array data were used to correct and authenticate several thousand gene models during the genome annotation process.

Animals↗

Bioinformatics-based identification of chemosensory proteins in African Malaria Mosquito, Anopheles gambiae.

Chemosensory proteins (CSPs) are identifiable by four spatially conserved Cysteine residues in their primary structure or by two disulfide bridges in their tertiary structure according to the previously identified olfactory specific-D related proteins. A genomics- and bioinformatics-based approach is taken in the present study to identify the putative CSPs in the malaria-carrying mosquito, Anopheles gambiae. The results show that five out of the nine annotated candidates are the most possible Anopheles CSPs of A. gambiae. This study lays the foundation for further functional identification of Anopheles CSPs, though all of these candidates need additional experimental verification.

Amino Acid Sequence↗

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics↗

Genome-wide analysis of mRNA lengths in Saccharomyces cerevisiae.

BACKGROUND: Although the protein-coding sequences in the Saccharomyces cerevisiae genome have been studied and annotated extensively, much less is known about the extent and characteristics of the untranslated regions of yeast mRNAs. RESULTS: We developed a 'Virtual Northern' method, using DNA microarrays for genome-wide systematic analysis of mRNA lengths. We used this method to measure mRNAs corresponding to 84% of the annotated open reading frames (ORFs) in the S. cerevisiae genome, with high precision and accuracy (measurement errors +/- 6-7%). We found a close linear relationship between mRNA lengths and the lengths of known or predicted translated sequences; mRNAs were typically around 300 nucleotides longer than the translated sequences. Analysis of genes deviating from that relationship identified ORFs with annotation errors, ORFs that appear not to be bona fide genes, and potentially novel genes. Interestingly, we found that systematic differences in the total length of the untranslated sequences in mRNAs were related to the functions of the encoded proteins. CONCLUSIONS: The Virtual Northern method provides a practical and efficient method for genome-scale analysis of transcript lengths. Approximately 12-15% of the yeast genome is represented in untranslated sequences of mRNAs. A systematic relationship between the lengths of the untranslated regions in yeast mRNAs and the functions of the proteins they encode may point to an important regulatory role for these sequences.

Blotting, Northern↗

RNAi-induced phenotypes suggest a novel role for a chemosensory protein CSP5 in the development of embryonic integument in the honeybee (Apis mellifera).

Small chemosensory proteins (CSPs) belong to a conserved, but poorly understood, protein family found in insects and other arthropods. They exhibit both broad and restricted expression patterns during development. In this paper, we used a combination of genome annotation, transcriptional profiling and RNA interference to unravel the functional significance of a honeybee gene (csp5) belonging to the CSP family. We show that csp5 expression resembles the maternal-zygotic pattern that is characterized by the initiation of transcription in the ovary and the replacement of maternal mRNA with embryonic mRNA. Blocking the embryonic expression of csp5 with double-stranded RNA causes abnormalities in all body parts where csp5 is highly expressed. The treated embryos show a "diffuse", often grotesque morphology, and the head skeleton appears to be severely affected. They are 'unable-to-hatch' and cannot progress to the larval stages. Our findings reveal a novel, essential role for this gene family and suggest that csp5 (unable-to-hatch) is an ectodermal gene involved in embryonic integument formation. Our study confirms the utility of an RNAi approach to functional characterization of novel developmental genes uncovered by the honeybee genome project and provides a starting point for further studies on embryonic integument formation in this insect.

Amino Acid Sequence↗

Helicobacter pylori flagellar hook-filament transition is controlled by a FliK functional homolog encoded by the gene HP0906.

Helicobacter pylori is a human gastric pathogen which is dependent on motility for infection. The H. pylori genome encodes a near-complete complement of flagellar proteins compared to model enteric bacteria. One of the few flagellar genes not annotated in H. pylori is that encoding FliK, a hook length control protein whose absence leads to a polyhook phenotype in Salmonella enterica. We investigated the role of the H. pylori gene HP0906 in flagellar biogenesis because of linkage to other flagellar genes, because of its transcriptional regulation pattern, and because of the properties of an ortholog in Campylobacter jejuni (N. Kamal and C. W. Penn, unpublished data). A nonpolar mutation of HP0906 in strain CCUG 17874 was generated by insertion of a chloramphenicol resistance marker. Cells of the mutant were almost completely nonmotile but produced sheathed, undulating polyhook structures at the cell pole. Expression of HP0906 in a Salmonella fliK mutant restored motility, confirming that HP0906 is the H. pylori fliK gene. Mutation of HP0906 caused a dramatic reduction in H. pylori flagellin protein production and a significant increase in production of the hook protein FlgE. The HP0906 mutant showed increased transcription of the flgE and flaB genes relative to the wild type, down-regulation of flaA transcription, and no significant change in transcription of the flagellar intermediate class genes flgM, fliD, and flhA. We conclude that the H. pylori HP0906 gene product is the hook length control protein FliK and that its function is required for turning off the sigma(54) regulon during progression of the flagellar gene expression cascade.

Amino Acid Sequence↗

Discovery of eight novel divergent homologs expressed in cattle placenta.

Ten divergent homologs were identified using a subtractive bioinformatic analysis of 12,614 cattle placenta expressed sequence tags followed by comparative, evolutionary, and gene expression studies. Among the 10 divergent homologs, 8 have not been identified previously. These were named as follows: cattle cerebrum and skeletal muscle-specific transcript 1 (CSSMST1), cattle intestine-specific transcript 1 (CIST1), hepatitis A virus cellular receptor 1 amino-terminal domain-containing protein (HAVCRNDP), prolactin-related proteins 8, 9, and 11 (PRP8, PRP9, and PRP11, respectively) and secreted and transmembrane protein 1A and 1B (SECTM1A and SECTM1B, respectively). In addition, two previously known divergent genes were identified, trophoblast Kunitz domain protein 1 (TKDP1) and a new splice variant of TKDP4. Nucleotide substitution analysis provided evidence for positive selection in members of the PRP gene family, SECTM1A and SECTM1B. Gene expression profiles, motif predictions, and annotations of homologous sequences indicate immunological and reproductive functions of the divergent homologs. The genes identified in this study are thus of evolutionary and physiological importance and may have a role in placental adaptations.

Amino Acid Sequence↗

The gene-protein database of Escherichia coli: edition 5.

The gene-protein database of Escherichia coli is both an index relating a gene to its protein product on two-dimensional gels, and a catalog of information about the function, regulation, and genetics of individual proteins obtained from two-dimensional gel analysis or collated from the literature. Edition 5 has 102 new entries--a 15% increase in the number of annotated two-dimensional gel spots. The large increase in this edition was accomplished in part by the use of a new method for expression analysis of ordered segments of the E. coli genome, which has resulted in linking 50 gel spots to their genes (or open reading frames) and another 45 to specific regions of the chromosome awaiting the availability of DNA sequence information. Communication of information from the scientific community resulted in additional identifications and regulatory information. To increase accessibility of the database it has been placed in the repository at the National Center for Biotechnology Information (NCBI) at the National Library of Medicine under the name ECO2DBASE. It will be updated twice yearly. This edition of the gene-protein database is estimated to contain entries for one-sixth of the protein-encoding genes of E. coli.

Bacterial Proteins↗

Hidden proteins encoded by non-canonical open reading frames: A review.

There is increasing evidence that translation is not limited to annotated protein-coding genes. Ribosome profiling sequencing, mass spectrometry-based proteomics, and immunopeptidomics have identified the productive translation of non-canonical open reading frames (ORFs). This suggests that the functional proteome includes not only conserved proteins but also proteins hidden in non-coding RNAs and de novo proteins. Some of these translated products are functional peptides, while others may be non-functional, potentially arising from evolutionary events. Several non-canonical ORF-encoded peptides have been found to regulate multiple physiological and pathological functions, particularly in cancer, immunity, and inflammation, indicating that they have potential as biomarkers and novel therapeutic targets. To better understand the diversity of functional peptides and translated non-canonical ORFs based on existing data, we summarize their classification according to transcriptional features and supporting evidence, including non-canonical ORFs located in ncRNAs and canonical mRNAs. This review provides a concise summary of the origin, discovery methods, and classification of non-canonical ORFs. It offers insights into the origins and functions of non-canonical ORF-encoded peptides from an evolutionary perspective, while also exploring the biological functions and regulatory mechanisms of these non-canonical ORF-encoded hidden proteins in tumorigenesis and progression.

Open Reading Frames↗

Monosomy for the most telomeric, gene-rich region of the short arm of human chromosome 16 causes minimal phenotypic effects.

We have examined the phenotypic effects of 21 independent deletions from the fully sequenced and annotated 356 kb telomeric region of the short arm of chromosome 16 (16p13.3). Fifteen genes contained within this region have been highly conserved throughout evolution and encode proteins involved in important housekeeping functions, synthesis of haemoglobin, signalling pathways and critical developmental pathways. Although a priori many of these genes would be considered candidates for critical haploinsufficient genes, none of the deletions within the 356 kb interval cause any discernible phenotype other than alpha thalassaemia whether inherited via the maternal or paternal line. These findings contrast with previous observations on patients with larger (> 1 Mb) deletions from the 16p telomere and therefore address the mechanisms by which monosomy gives rise to human genetic disease.

Adolescent↗

The subsystems approach to genome annotation and its use in the project to annotate 1000 genomes.

The release of the 1000th complete microbial genome will occur in the next two to three years. In anticipation of this milestone, the Fellowship for Interpretation of Genomes (FIG) launched the Project to Annotate 1000 Genomes. The project is built around the principle that the key to improved accuracy in high-throughput annotation technology is to have experts annotate single subsystems over the complete collection of genomes, rather than having an annotation expert attempt to annotate all of the genes in a single genome. Using the subsystems approach, all of the genes implementing the subsystem are analyzed by an expert in that subsystem. An annotation environment was created where populated subsystems are curated and projected to new genomes. A portable notion of a populated subsystem was defined, and tools developed for exchanging and curating these objects. Tools were also developed to resolve conflicts between populated subsystems. The SEED is the first annotation environment that supports this model of annotation. Here, we describe the subsystem approach, and offer the first release of our growing library of populated subsystems. The initial release of data includes 180 177 distinct proteins with 2133 distinct functional roles. This data comes from 173 subsystems and 383 different organisms.

Acyl Coenzyme A↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

Active components and potential mechanisms of Wuzhuyu decoction in the treatment of ethanol-induced acute gastric mucosal injury: a network pharmacology and experimental verification.

OBJECTIVE: To investigate the underlying mechanisms and active components of Wuzhuyu decoction (, WD) in alleviating ethanol-induced acute gastric mucosal injury (GMI) using an integrated approach of network pharmacology and experimental verification. METHODS: Sprague-Dawley rats were randomly divided into six groups: control (Con), model (Mod), bismuth potassium citrate (BPC), WD at low (WD-L), medium (WD-M), and high (WD-H) doses. Following seven days of continuous intragastric administration of the respective treatments, an ethanol-induced gastric mucosal injury model was established in all groups except the control group by oral gavage of anhydrous ethanol. The gastric mucosal injury index was evaluated, and pathological changes were assessed viahematoxylin and eosin (HE) staining. Levels of tumor necrosis factor-alpha (TNF-α), interleukin-1 beta (IL-1β), malondialdehyde (MDA), superoxide dismutase (SOD), and glutathione peroxidase (GSH-Px) were measured by enzyme-linked immunosorbent assay (ELISA). The chemical composition was identified by ultra-performance liquid chromatography-tandem mass spectrometry. Active compounds were screened using the Swiss-absorption, distribution, metabolism, and excretion database, and their potential targets were predicted using the Swiss Target Prediction database and bioinformatics annotation database for molecular mechanism. Simultaneously, disease targets related to GMI were retrieved from the online mendelian inheritance in man and GeneCards databases. A protein-protein interaction (PPI) network was constructed, and functional enrichment analyses of gene ontology (GO) and Kyoto encyclopedia of genes and genomes (KEGG) enrichment analyses were performed using the Metascape database. Key predictions from the network pharmacology analysis were subsequently verified through animal experiments. Protein expression levels of B-cell lymphoma-2 (Bcl-2), Bcl-2-associated X protein (Bax), Cleaved Caspase-3, and Cleaved Caspase-9 were analyzed by Western blot. Finally, molecular docking was performed using AutoDock Vina to investigate the interactions between the active components and core targets. RESULTS: WD treatment significantly reduced the gastric mucosal injury index and the levels of TNF-α, IL-1β, MDA, while it increased the activities of SOD and GSH-Px. Histopathological examination revealed marked improvement in gastric tissue morphology. A total of 145 compounds were identified in WD. Network pharmacology analysis identified 440 overlapping targets between WD and GMI. GO and KEGG enrichment analyses highlighted the apoptosis signaling pathway as a key mechanism for WD's protective effect against ethanol-induced GMI. Experimental validation demonstrated that WD treatment reduced the apoptosis of gastric mucosal epithelial cells, promoted the expression of Bcl-2, and inhibited the expression of Bax, Cleaved Caspase-3 and Cleaved Caspase-9. Molecular docking results indicated that dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin are potential active components in WD that contribute to the inhibition of apoptosis. CONCLUSIONS: WD alleviates ethanol-induced acute GMI, at least in part, by inhibiting the apoptosis. The primary active components responsible for this effect are dehydroevodiamine, rutaecarpine, evodiamine, hexahydrocurcumin, and isorhamnetin.

Drugs, Chinese Herbal↗

Phylogenomic analysis of the GIY-YIG nuclease superfamily.

BACKGROUND: The GIY-YIG domain was initially identified in homing endonucleases and later in other selfish mobile genetic elements (including restriction enzymes and non-LTR retrotransposons) and in enzymes involved in DNA repair and recombination. However, to date no systematic search for novel members of the GIY-YIG superfamily or comparative analysis of these enzymes has been reported. RESULTS: We carried out database searches to identify all members of known GIY-YIG nuclease families. Multiple sequence alignments together with predicted secondary structures of identified families were represented as Hidden Markov Models (HMM) and compared by the HHsearch method to the uncharacterized protein families gathered in the COG, KOG, and PFAM databases. This analysis allowed for extending the GIY-YIG superfamily to include members of COG3680 and a number of proteins not classified in COGs and to predict that these proteins may function as nucleases, potentially involved in DNA recombination and/or repair. Finally, all old and new members of the GIY-YIG superfamily were compared and analyzed to infer the phylogenetic tree. CONCLUSION: An evolutionary classification of the GIY-YIG superfamily is presented for the very first time, along with the structural annotation of all (sub)families. It provides a comprehensive picture of sequence-structure-function relationships in this superfamily of nucleases, which will help to design experiments to study the mechanism of action of known members (especially the uncharacterized ones) and will facilitate the prediction of function for the newly discovered ones.

Archaeal Proteins↗