Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The Protein Information Resource: an integrated public resource of functional annotation of proteins.

The Protein Information Resource (PIR) serves as an integrated public resource of functional annotation of protein data to support genomic/proteomic research and scientific discovery. The PIR, in collaboration with the Munich Information Center for Protein Sequences (MIPS) and the Japan International Protein Information Database (JIPID), produces the PIR-International Protein Sequence Database (PSD), the major annotated protein sequence database in the public domain, containing about 250 000 proteins. To improve protein annotation and the coverage of experimentally validated data, a bibliography submission system is developed for scientists to submit, categorize and retrieve literature information. Comprehensive protein information is available from iProClass, which includes family classification at the superfamily, domain and motif levels, structural and functional features of proteins, as well as cross-references to over 40 biological databases. To provide timely and comprehensive protein data with source attribution, we have introduced a non-redundant reference protein database, PIR-NREF. The database consists of about 800 000 proteins collected from PIR-PSD, SWISS-PROT, TrEMBL, GenPept, RefSeq and PDB, with composite protein names and literature data. To promote database interoperability, we provide XML data distribution and open database schema, and adopt common ontologies. The PIR web site (http://pir.georgetown.edu/) features data mining and sequence analysis tools for information retrieval and functional identification of proteins based on both sequence and annotation information. The PIR databases and other files are also available by FTP (ftp://nbrfa.georgetown.edu/pir_databases).

Amino Acid Sequence↗

AutoFACT: an automatic functional annotation and classification tool.

BACKGROUND: Assignment of function to new molecular sequence data is an essential step in genomics projects. The usual process involves similarity searches of a given sequence against one or more databases, an arduous process for large datasets. RESULTS: We present AutoFACT, a fully automated and customizable annotation tool that assigns biologically informative functions to a sequence. Key features of this tool are that it (1) analyzes nucleotide and protein sequence data; (2) determines the most informative functional description by combining multiple BLAST reports from several user-selected databases; (3) assigns putative metabolic pathways, functional classes, enzyme classes, GeneOntology terms and locus names; and (4) generates output in HTML, text and GFF formats for the user's convenience. We have compared AutoFACT to four well-established annotation pipelines. The error rate of functional annotation is estimated to be only between 1-2%. Comparison of AutoFACT to the traditional top-BLAST-hit annotation method shows that our procedure increases the number of functionally informative annotations by approximately 50%. CONCLUSION: AutoFACT will serve as a useful annotation tool for smaller sequencing groups lacking dedicated bioinformatics staff. It is implemented in PERL and runs on LINUX/UNIX platforms. AutoFACT is available at http://megasun.bch.umontreal.ca/Software/AutoFACT.htm.

Acanthamoeba castellanii↗

Automatic discovery of regulatory patterns in promoter regions based on whole cell expression data and functional annotation.

MOTIVATION: The whole genomes submitted to GenBank contain valuable information about the function of genes as well as the upstream sequences and whole cell expression provides valuable information on gene regulation. To utilize these large amounts of data for a biological understanding of the regulation of gene expression, new automatic methods for pattern finding are needed. RESULTS: Two word-analysis algorithms for automatic discovery of regulatory sequence elements have been developed. We show that sequence patterns correlated to whole cell expression data can be found using Kolmogorov-Smirnov tests on the raw data, thereby eliminating the need for clustering co-regulated genes. Regulatory elements have also been identified by systematic calculations of the significance of correlations between words found in the functional annotation of genes and DNA words occurring in their promoter regions. Application of these algorithms to the Saccharomyces cerevisiae genome and publicly available DNA array data sets revealed a highly conserved 9-mer occurring in the upstream regions of genes coding for proteasomal subunits. Several other putative and known regulatory elements were also found. AVAILABILITY: Upon request.

Algorithms↗

A novel method for automatic functional annotation of proteins.

MOTIVATION: To cope with the increasing amount of sequence data, reliable automatic annotation tools are required. The TrEMBL database contains together with SWISS-PROT nearly all publicly available protein sequences, but in contrast to SWISS-PROT only limited functional annotation. To improve this situation, we had to develop a method of automatic annotation that produces highly reliable functional prediction using the language and the syntax of SWISS-PROT. RESULTS: An algorithm was developed and successfully used for the automatic annotation of a testset of unknown proteins. The predicted information included description, function, catalytic activity, cofactors, pathway, subcellular location, quaternary structure, similarity to other protein, active sites, and keywords. The algorithm showed a low coverage (10%), but a high specificity and reliability. AVAILABILITY: The results can be obtained by anonymous ftp from ftp.ebi.ac.uk/pub/databases/sp_tr_nrdb. The source code is available on request from the authors.

Algorithms↗

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation↗

Functional annotation of class I lysyl-tRNA synthetase phylogeny indicates a limited role for gene transfer.

Functional and comparative genomic studies have previously shown that the essential protein lysyl-tRNA synthetase (LysRS) exists in two unrelated forms. Most prokaryotes and all eukaryotes contain a class II LysRS, whereas most archaea and a few bacteria contain a less common class I LysRS. In bacteria the class I LysRS is only found in the alpha-proteobacteria and a scattering of other groups, including the spirochetes, while the class I protein is by far the most common form of LysRS in archaea. To investigate this unusual distribution we functionally annotated a representative phylogenetic sampling of LysRS proteins. Class I LysRS proteins from a variety of bacteria and archaea were characterized in vitro by their ability to recognize Escherichia coli tRNA(Lys) anticodon mutants. Class I LysRS proteins were found to fall into two distinct groups, those that preferentially recognize the third anticodon nucleotide of tRNA(Lys) (U36) and those that recognize both the second and third positions (U35 and U36). Strong recognition of U35 and U36 was confined to the pyrococcus-spirochete grouping within the archaeal branch of the class I LysRS phylogenetic tree, while U36 recognition was seen in other archaea and an example from the alpha-proteobacteria. Together with the corresponding phylogenetic relationships, these results suggest that despite its comparative rarity the distribution of class I LysRS conforms to the canonical archaeal-bacterial division. The only exception, suggested from both functional and phylogenetic data, appears to be the horizontal transfer of class I LysRS from a pyrococcal progenitor to a limited number of bacteria.

Acylation↗

A functional annotation of subproteomes in human plasma.

The data collected by Human Proteome Organization's Plasma Proteome Pilot project phase was analyzed by members of our working group. Accordingly, a functional annotation of the human plasma proteome was carried out. Here, we report the findings of our analyses. First, bioinformatic analyses were undertaken to determine the likely sources of plasma proteins and to develop a protein interaction network of proteins identified in this project. Second, annotation of these proteins was performed in the context of functional subproteomes involved in the coagulation pathway, the mononuclear phagocytic system, the inflammation pathway, the cardiovascular system, and the liver; as well as the subset of proteins associated with DNA binding activities. Our analyses contributed to the Plasma Proteome Database (http://www.plasmaproteomedatabase.org), an annotated database of plasma proteins identified by HPPP as well as from other published studies. In addition, we address several methodological considerations including the selective enrichment of post-translationally modified proteins by the use of multi-lectin chromatography as well as the use of peptidomic techniques to characterize the low molecular weight proteins in plasma. Furthermore, we have performed additional analyses of peptide identification data to annotate cleavage of signal peptides, sites of intra-membrane proteolysis and post-translational modifications. The HPPP-organized, multi-laboratory effort, as described herein, resulted in much synergy and was essential to the success of this project.

Blood Coagulation↗

CASTp: computed atlas of surface topography of proteins with structural and topographical mapping of functionally annotated residues.

Cavities on a proteins surface as well as specific amino acid positioning within it create the physicochemical properties needed for a protein to perform its function. CASTp (http://cast.engr.uic.edu) is an online tool that locates and measures pockets and voids on 3D protein structures. This new version of CASTp includes annotated functional information of specific residues on the protein structure. The annotations are derived from the Protein Data Bank (PDB), Swiss-Prot, as well as Online Mendelian Inheritance in Man (OMIM), the latter contains information on the variant single nucleotide polymorphisms (SNPs) that are known to cause disease. These annotated residues are mapped to surface pockets, interior voids or other regions of the PDB structures. We use a semi-global pair-wise sequence alignment method to obtain sequence mapping between entries in Swiss-Prot, OMIM and entries in PDB. The updated CASTp web server can be used to study surface features, functional regions and specific roles of key residues of proteins.

Amino Acids↗

Ruleminer: a knowledge system for supporting high-throughput protein function annotations.

In this paper, we present RuleMiner, a knowledge system to facilitate a seamless integration of multi-sequence analysis tools and define profile-based rules for supporting high-throughput protein function annotations. This system consists of three essential components, Protein Function Groups (PFGs), PFG profiles and rules. The PFGs, established from an integrated analysis of current knowledge of protein functions from Swiss-Prot database and protein family-based sequence classifications, cover all possible cellular functions available in the database. The PFG profiles illustrate detailed protein features in the PFGs as in sequence conservations, the occurrences of sequence-based motifs, domains and species distributions. The rules, extracted from the PFG profiles, describe the clear relationships between these PFGs and all possible features. As a result, the RuleMiner is able to provide an enhanced capability for protein function analysis, such as results from the integrated sequence analysis tools for given proteins can be comparatively analyzed due to the clear feature-PFG relationships. Also, much needed guidance is readily available for such analysis. If the rules describe one-to-one (unique) relationships between the protein features and the PFGs, then these features can be utilized as unique functional identifiers and cellular functions of unknown proteins can be reliably determined. Otherwise, additional information has to be provided.

Algorithms↗

HapScope: a software system for automated and visual analysis of functionally annotated haplotypes.

We have developed a software analysis package, HapScope, which includes a comprehensive analysis pipeline and a sophisticated visualization tool for analyzing functionally annotated haplotypes. The HapScope analysis pipeline supports: (i) computational haplotype construction with an expectation-maximization or Bayesian statistical algorithm; (ii) SNP classification by protein coding change, homology to model organisms or putative regulatory regions; and (iii) minimum SNP subset selection by either a Brute Force Algorithm or a Greedy Partition Algorithm. The HapScope viewer displays genomic structure with haplotype information in an integrated environment, providing eight alternative views for assessing genetic and functional correlation. It has a user-friendly interface for: (i) haplotype block visualization; (ii) SNP subset selection; (iii) haplotype consolidation with subset SNP markers; (iv) incorporation of both experimentally determined haplotypes and computational results; and (v) data export for additional analysis. Comparison of haplotypes constructed by the statistical algorithms with those determined experimentally shows variation in haplotype prediction accuracies in genomic regions with different levels of nucleotide diversity. We have applied HapScope in analyzing haplotypes for candidate genes and genomic regions with extensive SNP and genotype data. We envision that the systematic approach of integrating functional genomic analysis with population haplotypes, supported by HapScope, will greatly facilitate current genetic disease research.

Algorithms↗

FunnyBase: a systems level functional annotation of Fundulus ESTs for the analysis of gene expression.

BACKGROUND: While studies of non-model organisms are critical for many research areas, such as evolution, development, and environmental biology, they present particular challenges for both experimental and computational genomic level research. Resources such as mass-produced microarrays and the computational tools linking these data to functional annotation at the system and pathway level are rarely available for non-model species. This type of "systems-level" analysis is critical to the understanding of patterns of gene expression that underlie biological processes. RESULTS: We describe a bioinformatics pipeline known as FunnyBase that has been used to store, annotate, and analyze 40,363 expressed sequence tags (ESTs) from the heart and liver of the fish, Fundulus heteroclitus. Primary annotations based on sequence similarity are linked to networks of systematic annotation in Gene Ontology (GO) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) and can be queried and computationally utilized in downstream analyses. Steps are taken to ensure that the annotation is self-consistent and that the structure of GO is used to identify higher level functions that may not be annotated directly. An integrated framework for cDNA library production, sequencing, quality control, expression data generation, and systems-level analysis is presented and utilized. In a case study, a set of genes, that had statistically significant regression between gene expression levels and environmental temperature along the Atlantic Coast, shows a statistically significant (P < 0.001) enrichment in genes associated with amine metabolism. CONCLUSION: The methods described have application for functional genomics studies, particularly among non-model organisms. The web interface for FunnyBase can be accessed at http://genomics.rsmas.miami.edu/funnybase/super_craw4/. Data and source code are available by request at jpaschall@bioinfobase.umkc.edu.

Animals↗

Functional annotation of the putative orphan Caenorhabditis elegans G-protein-coupled receptor C10C6.2 as a FLP15 peptide receptor.

This report describes the cloning and functional annotation of a Caenorhabditis elegans orphan G-protein-coupled receptor (GPCR) (C10C6.2) as a receptor for the FMRFamide-related peptides (FaRPs) encoded on the flp15 precursor gene, leading to the receptor designation FLP15-R. A cDNA encoding C10C6.2 was obtained using PCR techniques, confirmed identical to the Worm-pep-predicted sequence, and cloned into a vector appropriate for eucaryotic expression. A [35S]guanosine 5'-O-(thiotriphosphate) (GTPgammaS) assay with membranes prepared from Chinese hamster ovary (CHO) cells transiently transfected with FLP15-R was used as a read-out for receptor activation. FLP15-R was activated by putative FLP15 peptides, GGPQGPLRF-NH2 (FLP15-1), RGPSGPLRF-NH2 (FLP15-2A), its des-Arg1 counterpart, GPSGPLRF-NH2 (FLP15-2B), and to a lesser extent, by a tobacco hornworm Manduca sexta FaRP, GNSFLRFNH2 (F7G) (potency ranking FLP15-2A > FLP15-1 > FLP15-2B >> F7G). FLP15-R activation was abolished in the transfected cells pretreated with pertussis toxin, suggesting a preferential receptor coupling to Gi/Go proteins. The functional expression of FLP15-R in mammalian cells was temperature-dependent. Either no stimulation or significantly lower ligand-evoked [35S]GTPgammaS binding was observed in membranes prepared from transfected FLP15-R/CHO cells cultured at 37 degrees C. However, a 37 to 28 degrees C temperature shift implemented 24 h post-transfection consistently resulted in an improved activation signal and was essential for detectable functional expression of FLP15-R in CHO cells. To our knowledge, the FLP15 receptor is only the second deorphanized C. elegans neuropeptide GPCR reported to date.

Amino Acid Sequence↗

Comparing functional annotation analyses with Catmap.

BACKGROUND: Ranked gene lists from microarray experiments are usually analysed by assigning significance to predefined gene categories, e.g., based on functional annotations. Tools performing such analyses are often restricted to a category score based on a cutoff in the ranked list and a significance calculation based on random gene permutations as null hypothesis. RESULTS: We analysed three publicly available data sets, in each of which samples were divided in two classes and genes ranked according to their correlation to class labels. We developed a program, Catmap (available for download at http://bioinfo.thep.lu.se/Catmap), to compare different scores and null hypotheses in gene category analysis, using Gene Ontology annotations for category definition. When a cutoff-based score was used, results depended strongly on the choice of cutoff, introducing an arbitrariness in the analysis. Comparing results using random gene permutations and random sample permutations, respectively, we found that the assigned significance of a category depended strongly on the choice of null hypothesis. Compared to sample label permutations, gene permutations gave much smaller p-values for large categories with many coexpressed genes. CONCLUSIONS: In gene category analyses of ranked gene lists, a cutoff independent score is preferable. The choice of null hypothesis is very important; random gene permutations does not work well as an approximation to sample label permutations.

Classification↗

Knowledge-based voting algorithm for automated protein functional annotation.

Automated annotation of high-throughput genome sequences is one of the earliest steps toward a comprehensive understanding of the dynamic behavior of living organisms. However, the step is often error-prone because of its underlying algorithms, which rely mainly on a simple similarity analysis, and lack of guidance from biological rules. We present herein a knowledge-based protein annotation algorithm. Our objectives are to reduce errors and to improve annotation confidences. This algorithm consists of two major components: a knowledge system, called "RuleMiner," and a voting procedure. The knowledge system, which includes biological rules and functional profiles for each function, provides a platform for seamless integration of multiple sequence analysis tools and guidance for function annotation. The voting procedure, which relies on the knowledge system, is designed to make (possibly) unbiased judgments in functional assignments among complicated, sometimes conflicting, information. We have applied this algorithm to 10 prokaryotic bacterial genomes and observed a significant improvement in annotation confidences. We also discuss the current limitations of the algorithm and the potential for future improvement.

Algorithms↗

Functional annotation of the Arabidopsis P450 superfamily based on large-scale co-expression analysis.

Cytochrome P450 mono-oxygenases play prominent roles in a diverse set of metabolic pathways, but the function of most of these enzymes remains obscure. A bottleneck in the functional genomics of this superfamily constitutes hypothesis generation to identify potential substrates (or substrate classes) individual P450s may act on. We used publicly available large-scale expression data to perform co-expression analysis comparing the expression matrix of each P450 with those from more than 4000 selected genes across thousands of microarrays. Based on functional annotations of co-expressed genes from a diverse set of databases, co-expressed pathways were thus identified for each P450. Using this approach, most P450s with known functions were placed into their respective pathways, thereby proofing the concept. As examples, pathway mapping results identifying novel P450s potentially acting on flower-specific monoterpenes and root-specific triterpenes are described. Co-expression results for all Arabidopsis P450s will be presented as a web resource on the 'CYPedia' web pages (http://ibmp.u-strasbg.fr/CYPedia/).

Arabidopsis↗

Genomic functional annotation using co-evolution profiles of gene clusters.

BACKGROUND: The current speed of sequencing already exceeds the capability of annotation, creating a potential bottleneck. A large proportion of the genes in microbial genomes remains uncharacterized. Here we propose a new method for functional annotation using the conservation patterns of gene clusters. If several gene clusters show the same coevolution pattern across different genomes it is reasonable to infer they are functionally related. The gene cluster phylogenetic profile integrates chromosomal proximity information and phylogenetic profile information and allows us to infer functional dependences between the gene clusters even at great distance on the chromosome. RESULTS: As a proof of concept, we applied our method to the genome of Escherichia coli K12 strain. Our method establishes functional relationships among 176 gene clusters, comprising 738 E. coli genes. The accuracy of pair phylogenetic profiles was compared with the single-gene phylogenetic profile and was shown to be higher. As a result, we are able to suggest functional roles for several previously unknown genes or unknown genomic regions in E. coli. We also examined the robustness of coevolution signals across a larger set of genomes and suggest a possible upper limit of accuracy for the phylogenetic profile methods. CONCLUSIONS: The higher-order phylogenetic profiles, such as the gene-pair phylogenetic profiles, can detect functional dependences that are missed by using conventional single-gene phylogenetic profile or the chromosomal proximity method only. We show that the gene-pair phylogenetic profile is more accurate than the single-gene phylogenetic profiles.

Alleles↗

Gene ontology application to genomic functional annotation, statistical analysis and knowledge mining.

While a massive amount of biomolecular information is increasingly accumulating in different databanks, on the other hand high-throughput technologies are generating a great quantity of data that need to be annotated with the genomic information available, and interpreted. To this aim, the use of specific ontologies can greatly help either in integrating different information stored within heterogeneous databanks, or in identifying and clustering sequence data sharing common characteristics. In the molecular biology domain, the Gene Ontology (GO) is the most developed and widely used ontology. To demonstrate its great utility in the annotation and biological interpretation of gene sets obtained by means of high-throughput experiments, we implemented the web application here described. It enables functional annotations of a given gene set on a genomic scale and across different species. Within our application the annotations provided by the GO vocabulary allow either to easily bind several information from different resources, or to cluster annotated genes according to their biological characteristics. Through the GO structure it is also possible to represent biological concepts with different specificity levels, from very general to very precise concepts. Furthermore, the statistical evaluation of the categorizations provided by the GO annotations enables to highlight the most significant biological characteristics of a gene set, and therefore to mine knowledge from data. Our created tool meets the need to manage a vast quantity of biological data with a simple user interface adapt also for users with limited informatics knowledge, leading them to evaluate the functional significance of experiment's results with graphical views and statistical indexes in a well-known web browser user interface.

Genomics↗

Multiple-sequence functional annotation and the generalized hidden Markov phylogeny.

MOTIVATION: Phylogenetic shadowing is a comparative genomics principle that allows for the discovery of conserved regions in sequences from multiple closely related organisms. We develop a formal probabilistic framework for combining phylogenetic shadowing with feature-based functional annotation methods. The resulting model, a generalized hidden Markov phylogeny (GHMP), applies to a variety of situations where functional regions are to be inferred from evolutionary constraints. RESULTS: We show how GHMPs can be used to predict complete shared gene structures in multiple primate sequences. We also describe shadower, our implementation of such a prediction system. We find that shadower outperforms previously reported ab initio gene finders, including comparative human-mouse approaches, on a small sample of diverse exonic regions. Finally, we report on an empirical analysis of shadower's performance which reveals that as few as five well-chosen species may suffice to attain maximal sensitivity and specificity in exon demarcation. AVAILABILITY: A Web server is available at http://bonaire.lbl.gov/shadower

Algorithms↗