Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals↗

Phydbac "Gene Function Predictor": a gene annotation tool based on genomic context analysis.

BACKGROUND: The large amount of completely sequenced genomes allows genomic context analysis to predict reliable functional associations between prokaryotic proteins. Major methods rely on the fact that genes encoding physically interacting partners or members of shared metabolic pathways tend to be proximate on the genome, to evolve in a correlated manner and to be fused as a single sequence in another organism. RESULTS: The new "Gene Function Predictor", linked to the web server Phydbac proposes putative associations between Escherichia coli K-12 proteins derived from a combination of these methods. We show that associations made by this tool are more accurate than linkages found in the other established databases. Predicted assignments to GO categories, based on pre-existing functional annotations of associated proteins are also available. This new database currently holds 9,379 pairwise links at an expected success rate of at least 80%, the 6,466 functional predictions to GO terms derived from these links having a level of accuracy higher than 70%. CONCLUSION: The "Gene Function Predictor" is an automatic tool that aims to help biologists by providing them hypothetical functional predictions out of genomic context characteristics. The "Gene Function predictor" is available at http://www.igs.cnrs-mrs.fr/phydbac/indexPS.html.

Algorithms↗

In vivo functional proteomics: mammalian genome annotation using CD-tagging.

A self-inactivating CD-tagging retroviral vector was used to introduce epitope and GFP tags into genes and proteins in NIH 3T3 cells. Several hundred cell clones, each expressing GFP fluorescence in a distinctive pattern, were isolated. Molecular analysis showed that a wide variety of genes and proteins, some known and some newly discovered, had been tagged. The analysis also revealed that, in the great majority of instances, the abundance and cellular location of the tagged protein mirrored that of its untagged counterpart. This approach provides a systematic means for the functional annotation of mammalian genomes and proteomes in living cells.

3T3 Cells↗

Microarray annotation and biological information on function.

OBJECTIVES: Many methods for statistical analysis of gene expression studies by DNA microarrays produce lists of genes as output. To understand gene lists in terms of traditional biology, e.g. which pathways may be affected, it is necessary to get appropriate annotations for the probes on an array. METHODS: Problems arise with the different sources that have been used by manufacturers to design microarray probes, and their association to biological entities like genes, transcripts and proteins. Function annotation is of crucial importance, and systems like Gene Ontology can be used for this purpose. It arranges annotation terms in a hierarchical manner and thus makes annotations in a gene list amenable to automated analysis. RESULTS: Several methods for analyses of gene function are described. The hierarchical nature of systems like Gene Ontology particularly suggests using methods from graph theory. CONCLUSIONS: The main problem in annotating microarray probes and inferring affected functional modules is the incompleteness and degree of error in current biological databases. Initial approaches to make use of functional annotation exist, but have to be extended, in particular with respect to estimating the statistical significance of results.

Computational Biology↗

Learning statistical models for annotating proteins with function information using biomedical text.

BACKGROUND: The BioCreative text mining evaluation investigated the application of text mining methods to the task of automatically extracting information from text in biomedical research articles. We participated in Task 2 of the evaluation. For this task, we built a system to automatically annotate a given protein with codes from the Gene Ontology (GO) using the text of an article from the biomedical literature as evidence. METHODS: Our system relies on simple statistical analyses of the full text article provided. We learn n-gram models for each GO code using statistical methods and use these models to hypothesize annotations. We also learn a set of Naïve Bayes models that identify textual clues of possible connections between the given protein and a hypothesized annotation. These models are used to filter and rank the predictions of the n-gram models. RESULTS: We report experiments evaluating the utility of various components of our system on a set of data held out during development, and experiments evaluating the utility of external data sources that we used to learn our models. Finally, we report our evaluation results from the BioCreative organizers. CONCLUSION: We observe that, on the test data, our system performs quite well relative to the other systems submitted to the evaluation. From other experiments on the held-out data, we observe that (i) the Naïve Bayes models were effective in filtering and ranking the initially hypothesized annotations, and (ii) our learned models were significantly more accurate when external data sources were used during learning.

Bayes Theorem↗

The Annotated Blueprint: Integrated Functional Genomic Resources for a model Tetraploid Wheat Triticum turgidum cv. Kronos.

Triticum turgidum cv. Kronos is a tetraploid wheat cultivar that underpins one of the richest community platforms for functional genomics. Over the past decade, about 3,000 exome- and promoter-capture datasets, linked to mutagenized seed stocks, and transcriptomic and phenotypic resources have accumulated, yet the absence of a reference genome has constrained their impact. Here, we present a chromosome-scale reference genome of Kronos with high-confidence annotations, including manual curation of over 1,000 disease resistance (NLR) genes. This reference revealed previously hidden NLR diversity and clarified their genomic organization at chromosomal ends. Re-analysis of exome- and promoter-capture datasets enabled high-resolution mutation discovery in genes and regulatory regions that were previously inaccessible, uncovering the full standing variation present in Kronos mutant lines. We further re-curated transcriptomic and small RNA datasets, generating improved, genome-wide maps of microRNAs and phasiRNAs important for wheat development. Collectively, these resources elevate Kronos to reference quality and establish it as a versatile platform for functional and translational wheat research.

Journal Article↗

Functional analysis and annotation of the virulence plasmid pMUM001 from Mycobacterium ulcerans.

The presence of a 174 kb plasmid called pMUM001 in Mycobacterium ulcerans, the first example of a mycobacterial plasmid encoding a virulence determinant, was recently reported. Over half of pMUM001 is devoted to six genes, three of which encode giant polyketide synthases (PKS) that produce mycolactone, an unusual cytotoxic lipid produced by M. ulcerans. In this present study the remaining 75 non-PKS-associated protein-coding sequences (CDS) are analysed and it is shown that pMUM001 is a low-copy-number element with a functional ori that supports replication in Mycobacterium marinum but not in the fast-growing mycobacteria Mycobacterium smegmatis and Mycobacterium fortuitum. Sequence analyses revealed a highly mosaic plasmid gene structure that is reminiscent of other large plasmids. Insertion sequences (IS) and fragments of IS, some previously unreported, are interspersed among functional gene clusters, such as those genes involved in plasmid replication, the synthesis of mycolactone, and a potential phosphorelay signal transduction system. Among the IS present on pMUM001 were multiple copies of the high-copy-number M. ulcerans elements IS2404 and IS2606. No plasmid transfer systems were identified, suggesting that trans-acting factors are required for mobilization. The results presented here provide important insights into this unusual virulence plasmid from an emerging but neglected human pathogen.

Bacterial Proteins↗

Automatic annotation of protein motif function with Gene Ontology terms.

BACKGROUND: Conserved protein sequence motifs are short stretches of amino acid sequence patterns that potentially encode the function of proteins. Several sequence pattern searching algorithms and programs exist foridentifying candidate protein motifs at the whole genome level. However, a much needed and important task is to determine the functions of the newly identified protein motifs. The Gene Ontology (GO) project is an endeavor to annotate the function of genes or protein sequences with terms from a dynamic, controlled vocabulary and these annotations serve well as a knowledge base. RESULTS: This paper presents methods to mine the GO knowledge base and use the association between the GO terms assigned to a sequence and the motifs matched by the same sequence as evidence for predicting the functions of novel protein motifs automatically. The task of assigning GO terms to protein motifs is viewed as both a binary classification and information retrieval problem, where PROSITE motifs are used as samples for mode training and functional prediction. The mutual information of a motif and aGO term association is found to be a very useful feature. We take advantage of the known motifs to train a logistic regression classifier, which allows us to combine mutual information with other frequency-based features and obtain a probability of correct association. The trained logistic regression model has intuitively meaningful and logically plausible parameter values, and performs very well empirically according to our evaluation criteria. CONCLUSIONS: In this research, different methods for automatic annotation of protein motifs have been investigated. Empirical result demonstrated that the methods have a great potential for detecting and augmenting information about the functions of newly discovered candidate protein motifs.

Amino Acid Motifs↗

Differential detergent fractionation for non-electrophoretic eukaryote cell proteomics.

Differential detergent fractionation (DDF), which relies on detergents to sequentially extract proteins from eukaryotic cells, has been used to increase proteome coverage of 2D-PAGE. Here, we used DDF extraction in conjunction with the nonelectrophoretic proteomics method of liquid chromatography and electrospray ionization tandem mass spectrometry. We demonstrate that DDF can be used with 2D-LC ESI MS2 for comprehensive cellular proteomics, including a large proportion of membrane proteins. Compared to some published methods designed to isolate membrane proteins specifically, DDF extraction yields comprehensive proteomes which include twice as many membrane proteins. Two-thirds of these membrane proteins have more than one trans-membrane domain. Since DDF separates proteins based upon their physicochemistry and subcellular localization, this method also provides data useful for functional genome annotation. As more genome sequences are completed, methods which can aid in functional annotation will become increasingly important.

Animals↗

Functional genomics by integrated analysis of metabolome and transcriptome of Arabidopsis plants over-expressing an MYB transcription factor.

The integration of metabolomics and transcriptomics can provide precise information on gene-to-metabolite networks for identifying the function of unknown genes unless there has been a post-transcriptional modification. Here, we report a comprehensive analysis of the metabolome and transcriptome of Arabidopsis thaliana over-expressing the PAP1 gene encoding an MYB transcription factor, for the identification of novel gene functions involved in flavonoid biosynthesis. For metabolome analysis, we performed flavonoid-targeted analysis by high-performance liquid chromatography-mass spectrometry and non-targeted analysis by Fourier-transform ion-cyclotron mass spectrometry with an ultrahigh-resolution capacity. This combined analysis revealed the specific accumulation of cyanidin and quercetin derivatives, and identified eight novel anthocyanins from an array of putative 1800 metabolites in PAP1 over-expressing plants. The transcriptome analysis of 22,810 genes on a DNA microarray revealed the induction of 38 genes by ectopic PAP1 over-expression. In addition to well-known genes involved in anthocyanin production, several genes with unidentified functions or annotated with putative functions, encoding putative glycosyltransferase, acyltransferase, glutathione S-transferase, sugar transporters and transcription factors, were induced by PAP1. Two putative glycosyltransferase genes (At5g17050 and At4g14090) induced by PAP1 expression were confirmed to encode flavonoid 3-O-glucosyltransferase and anthocyanin 5-O-glucosyltransferase, respectively, from the enzymatic activity of their recombinant proteins in vitro and results of the analysis of anthocyanins in the respective T-DNA-inserted mutants. The functional genomics approach through the integration of metabolomics and transcriptomics presented here provides an innovative means of identifying novel gene functions involved in plant metabolism.

Arabidopsis↗

Annotating enzymes of unknown function: N-formimino-L-glutamate deiminase is a member of the amidohydrolase superfamily.

The functional assignment of enzymes that catalyze unknown chemical transformations is a difficult problem. The protein Pa5106 from Pseudomonas aeruginosa has been identified as a member of the amidohydrolase superfamily by a comprehensive amino acid sequence comparison with structurally authenticated members of this superfamily. The function of Pa5106 has been annotated as a probablechlorohydrolase or cytosine deaminase. A close examination of the genomic content of P. aeruginosa reveals that the gene for this protein is in close proximity to genes included in the histidine degradation pathway. The first three steps for the degradation of histidine include the action of HutH, HutU, and HutI to convert L-histidine to N-formimino-L-glutamate. The degradation of N-formimino-L-glutamate to L-glutamate can occur by three different pathways. Three proteins in P. aeruginosa have been identified that catalyze two of the three possible pathways for the degradation of N-formimino-L-glutamate. The protein Pa5106 was shown to catalyze the deimination of N-formimino-L-glutamate to ammonia and N-formyl-L-glutamate, while Pa5091 catalyzed the hydrolysis of N-formyl-L-glutamate to formate and L-glutamate. The protein Pa3175 is dislocated from the hut operon and was shown to catalyze the hydrolysis of N-formimino-L-glutamate to formamide and L-glutamate. The reason for the coexistence of two alternative pathways for the degradation of N-formimino-L-glutamate in P. aeruginosa is unknown.

Amidohydrolases↗

PIR: a new resource for bioinformatics.

UNLABELLED: The Protein Information Resource (PIR) has greatly expanded its Web site and developed a set of interactive search and analysis tools to facilitate the analysis, annotation, and functional identification of proteins. New search engines have been implemented to combine sequence similarity search results with database annotation information. The new PIR search systems have proved very useful in providing enriched functional annotation of protein sequences, determining protein superfamily-domain relationships, and detecting annotation errors in genomic database archives. AVAILABILITY: http://pir.georgetown.edu/. CONTACT: mcgarvey@nbrf.georgetown.edu

Animals↗

GenomeRNAi: a database for cell-based RNAi phenotypes.

RNA interference (RNAi) has emerged as a powerful tool to generate loss-of-function phenotypes in a variety of organisms. Combined with the sequence information of almost completely annotated genomes, RNAi technologies have opened new avenues to conduct systematic genetic screens for every annotated gene in the genome. As increasing large datasets of RNAi-induced phenotypes become available, an important challenge remains the systematic integration and annotation of functional information. Genome-wide RNAi screens have been performed both in Caenorhabditis elegans and Drosophila for a variety of phenotypes and several RNAi libraries have become available to assess phenotypes for almost every gene in the genome. These screens were performed using different types of assays from visible phenotypes to focused transcriptional readouts and provide a rich data source for functional annotation across different species. The GenomeRNAi database provides access to published RNAi phenotypes obtained from cell-based screens and maps them to their genomic locus, including possible non-specific regions. The database also gives access to sequence information of RNAi probes used in various screens. It can be searched by phenotype, by gene, by RNAi probe or by sequence and is accessible at http://rnai.dkfz.de.

Animals↗

GFINDer: Genome Function INtegrated Discoverer through dynamic annotation, statistical analysis, and mining.

Statistical and clustering analyses of gene expression results from high-density microarray experiments produce lists of hundreds of genes regulated differentially, or with particular expression profiles, in the conditions under study. Independent of the microarray platforms and analysis methods used, these lists must be biologically interpreted to gain a better knowledge of the patho-physiological phenomena involved. To this end, numerous biological annotations are available within heterogeneous and widely distributed databases. Although several tools have been developed for annotating lists of genes, most of them do not give methods for evaluating the relevance of the annotations provided, or for estimating the functional bias introduced by the gene set on the array used to identify the gene list considered. We developed Genome Functional INtegrated Discoverer (GFINDer), a web server able to automatically provide large-scale lists of user-classified genes with functional profiles biologically characterizing the different gene classes in the list. GFINDer automatically retrieves annotations of several functional categories from different sources, identifies the categories enriched in each class of a user-classified gene list and calculates statistical significance values for each category. Moreover, GFINDer enables the functional classification of genes according to mined functional categories and the statistical analysis is of the classifications obtained, aiding better interpretation of microarray experiment results. GFINDer is available online at http://www.medinfopoli.polimi.it/GFINDer/.

Computational Biology↗

Molecular Characterization of Listeria monocytogenes Isolated from Retail Yak Meat in Nyingchi, Xizang, China.

Listeria monocytogenes is a Gram-positive zoonotic pathogen responsible for listeriosis, a severe foodborne disease with high mortality in humans and animals. This study aimed to investigate the molecular epidemiology and genomic characteristics of L. monocytogenes isolated from raw yak meat in Nyingchi, Xizang, China. A total of 231 yak-related samples were collected in Nyingchi, consisting of 214 retail raw yak meat samples, 14 farm environmental samples, and 3 nearby water source samples. L. monocytogenes isolates were identified and characterized using culture-based methods, PCR serotyping, and whole-genome sequencing (WGS). Bioinformatic analyses were performed for virulence, antimicrobial resistance, and functional gene annotation using KEGG and COG databases. The overall contamination rate of Lm was 13.08% (28/214) for retail raw yak meat samples, whereas no isolates were recovered from 14 farm environmental samples (0.00%, 0/14) and 3 nearby water source samples (0.00%, 0/3). The serotypes of isolates were 1/2a (9/28, 32.14%), 1/2b (7/28, 25.00%), and 1/2c (12/28, 42.86%). These 28 isolates exhibited varied antimicrobial resistance profiles, with universal resistance to trimethoprim-sulfamethoxazole, high resistance to erythromycin and clindamycin, and low resistance to vancomycin. MLST analysis revealed seven sequence types (STs): ST9 (12/28, 42.86%), ST619 (6/28, 21.43%), ST8 (6/28, 21.43%), ST7 (1/28, 3.57%), ST87 (1/28, 3.57%), ST121 (1/28, 3.57%), ST297 (1/28, 3.57%). ST619 isolates harbored multiple virulence genes, including those located on Listeria pathogenicity islands LIPI-1, LIPI-3, and LIPI-4, indicating high genomic potential for virulence. Representative isolate Y2 (ST619) possessed a 3,009,858 bp genome with 3036 coding genes, four genomic islands, and two prophages. Functional annotation revealed enrichment of genes involved in carbohydrate transport and metabolism and amino acid biosynthesis pathways. Our findings provide the first genomic insight into L. monocytogenes contamination in yak meat from Nyingchi, Xizang, China, highlighting the urgent need to strengthen food safety monitoring and hygiene management in this region.

Listeria monocytogenes↗

EST sequencing and time course microarray hybridizations identify more than 700 Medicago truncatula genes with developmental expression regulation in flowers and pods.

To evaluate the molecular mechanisms during pod and seed formation in legumes, starting with the development of reproductive organs, we constructed two cDNA libraries from developing flowers (MtFLOW) and pods including seeds (MtPOSE) of the model plant Medicago truncatula Gaertner. A total of 2,516 expressed sequence tags (ESTs) clustered into 1,776 nonredundant sequences (2k-set), which were annotated and assigned to functional classes. While about 30% of the ESTs encoded proteins of yet unknown function, typical annotations pointed to seed storage proteins, LTPs and lipoxygenases. The 2k-set was used to upgrade Mt6k-RIT microarrays (Küster et al. in J Biotechnol 108: 95, 2004) to Mt8k versions representing approximately 6,300 nonredundant M. truncatula genes. These were used to perform time course expression profiling studies based on hybridizations of samples that covered eight different developmental stages from flower buds to almost mature pods versus leaves as a common reference. About 180 up- and 70 downregulated genes were typically found for each stage and in total, 782 genes were either twofold up- or downregulated in at least one of the eight stages investigated. Based on this set, a combination of self-organizing map and hierarchical clustering revealed genes displaying expression regulation during characteristic stages of M. truncatula flower and pod development. Amongst those, several genes encoded proteins related to seed metabolism and development including novel regulators and proteins involved in signaling.

Expressed Sequence Tags↗

Novel targets of ANG II regulation in mouse heart identified by serial analysis of gene expression.

Although the central role of ANG II in cardiovascular homeostasis is well appreciated, the molecular circuitry of its many actions is not completely understood. With the use of serial analysis of gene expression to assess global transcriptional changes in the heart of mice after continuous 7-day ANG II administration, we identified patterns of gene expression indicative of cardiac remodeling, including coordinate regulation of genes previously described in a context of processes associated with hypertrophy and fibrosis. In addition, we discovered several novel ANG II targets, including characterized genes of known function, recently annotated genes of unknown function, and the putative genes not yet present in current databases. The serial analysis of gene expression approach to assess the role of ANG II presented in this report provides new venues for inquiries into ANG II-mediated cardiac function.

Angiotensin II↗