Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The CATH protein family database: a resource for structural and functional annotation of genomes.

Over the last decade, there have been huge increases in the numbers of protein sequences and structures determined. In parallel, many methods have been developed for recognising similarities between these proteins, arising from their common evolutionary background, and for clustering such relatives into protein families. Here we review some of the protein family resources available to the biologist and describe how these can be used to provide structural and functional annotations for newly determined sequences. In particular we describe recent developments to the CATH domain database of protein structural families which have facilitated genome annotation and which have also revealed important caveats that must be considered when transferring functional data between homologous proteins.

Databases, Protein↗

Metalloproteomics: high-throughput structural and functional annotation of proteins in structural genomics.

A high-throughput method for measuring transition metal content based on quantitation of X-ray fluorescence signals was used to analyze 654 proteins selected as targets by the New York Structural GenomiX Research Consortium. Over 10% showed the presence of transition metal atoms in stoichiometric amounts; these totals as well as the abundance distribution are similar to those of the Protein Data Bank. Bioinformatics analysis of the identified metalloproteins in most cases supported the metalloprotein annotation; identification of the conserved metal binding motif was also shown to be useful in verifying structural models of the proteins. Metalloproteomics provides a rapid structural and functional annotation for these sequences and is shown to be approximately 95% accurate in predicting the presence or absence of stoichiometric metal content. The project's goal is to assay at least 1 member from each Pfam family; approximately 500 Pfam families have been characterized with respect to transition metal content so far.

Binding Sites↗

Bioinformatics in reproductive biology--functional annotation based on comparative sequence analysis.

Recent studies of the genomes of a variety of model organisms have provided an unprecedented opportunity to identify and characterize all signaling molecules in the human genome. Regardless of the approaches used to decipher gene characteristics and their role in physiology, pairwise sequence comparison represents the fundamental bioinformatic tool for initial functional annotation of newly identified genes. Because genes evolved from duplication and adapted to different evolutionary niches for each organism during speciation, detailed sequence analysis could provide additional information on the biochemical and biological characteristics of novel genes. In addition, the integration of sequence-based gene discovery with phylogeny-based function prediction leads to a more complete understanding of the signaling pathways. For example, detailed pairwise sequence analysis has led to the identification of (1) stresscopin (SCP) and stresscopin-related peptides (SRP) as the CRH-related genes, (2) multiple relaxin-like factor genes, and (3) the novel glycoprotein hormone subunit family genes, alpha2 and beta5. Furthermore, based on the understanding that ligands and receptors coevolved during evolution, we have identified a variety of novel extracellular signaling polypeptides including (1) stresscopin and stresscopin-related peptides as selective ligands for the type 2 CRH receptor, (2) the pregnancy hormone, relaxin, and related peptides that activate two orphan G protein-coupled receptors (GPCRs), LGR7 and LGR8, and (3) alpha2 and beta5 that form a heterodimer capable of activating the TSH receptor. Thus, detailed studies on the characteristics and evolution of gene sequences have provided an inroad to the elucidation of novel signaling polypeptides and the associated signal transduction pathways.

Animals↗

Functional annotation of a full-length mouse cDNA collection.

The RIKEN Mouse Gene Encyclopaedia Project, a systematic approach to determining the full coding potential of the mouse genome, involves collection and sequencing of full-length complementary DNAs and physical mapping of the corresponding genes to the mouse genome. We organized an international functional annotation meeting (FANTOM) to annotate the first 21,076 cDNAs to be analysed in this project. Here we describe the first RIKEN clone collection, which is one of the largest described for any organism. Analysis of these cDNAs extends known gene families and identifies new ones.

Animals↗

Global protein function annotation through mining genome-scale data in yeast Saccharomyces cerevisiae.

As we are moving into the post genome-sequencing era, various high-throughput experimental techniques have been developed to characterize biological systems on the genomic scale. Discovering new biological knowledge from the high-throughput biological data is a major challenge to bioinformatics today. To address this challenge, we developed a Bayesian statistical method together with Boltzmann machine and simulated annealing for protein functional annotation in the yeast Saccharomyces cerevisiae through integrating various high-throughput biological data, including yeast two-hybrid data, protein complexes and microarray gene expression profiles. In our approach, we quantified the relationship between functional similarity and high-throughput data, and coded the relationship into 'functional linkage graph', where each node represents one protein and the weight of each edge is characterized by the Bayesian probability of function similarity between two proteins. We also integrated the evolution information and protein subcellular localization information into the prediction. Based on our method, 1802 out of 2280 unannotated proteins in yeast were assigned functions systematically.

Bayes Theorem↗

Functional annotation of deubiquitinating enzymes using RNA interference.

Protein ubiquitination is a dynamic process, depending on a tightly regulated balance between the activity of ubiquitin ligases and their antagonists, the ubiquitin-specific proteases or deubiquitinating enzymes. The family of ubiquitin ligases has been studied intensively and it is well established that their deregulation contributes to diverse disease processes, including cancer. Much less is known about the function and regulation of the large group of deubiquitinating enzymes. This chapter describes how RNA interference against deubiquitinating enzymes can be used to elucidate their function. The application of this technology will greatly improve the functional annotation of this family of proteases.

Cell Line, Tumor↗

Recent developments in microarray-based enzyme assays: from functional annotation to substrate/inhibitor fingerprinting.

Recent advances in proteomics have provided impetus towards the development of robust technologies for high-throughput studies of enzymes. The term "catalomics" defines an emerging '-omics' field in which high-throughput studies of enzymes are carried out by using advanced chemical proteomics approaches. Of the various available methods, microarrays have emerged as a powerful and versatile platform to accelerate not only the functional annotation but also the substrate and inhibitor specificity (e.g. substrate and inhibitor fingerprinting, respectively) of enzymes. Herein, we review recent developments in the fabrication of various types of microarray technologies (protein-, peptide- and small-molecule-based microarrays) and their applications in high-throughput characterizations of enzymes.

Animals↗

BioBuilder as a database development and functional annotation platform for proteins.

BACKGROUND: The explosion in biological information creates the need for databases that are easy to develop, easy to maintain and can be easily manipulated by annotators who are most likely to be biologists. However, deployment of scalable and extensible databases is not an easy task and generally requires substantial expertise in database development. RESULTS: BioBuilder is a Zope-based software tool that was developed to facilitate intuitive creation of protein databases. Protein data can be entered and annotated through web forms along with the flexibility to add customized annotation features to protein entries. A built-in review system permits a global team of scientists to coordinate their annotation efforts. We have already used BioBuilder to develop Human Protein Reference Database http://www.hprd.org, a comprehensive annotated repository of the human proteome. The data can be exported in the extensible markup language (XML) format, which is rapidly becoming as the standard format for data exchange. CONCLUSIONS: As the proteomic data for several organisms begins to accumulate, BioBuilder will prove to be an invaluable platform for functional annotation and development of customizable protein centric databases. BioBuilder is open source and is available under the terms of LGPL.

Computational Biology↗

Correlations between causal effect sizes of proximal SNPs vary with functional annotations and implicate stabilizing selection.

Causal disease effect sizes of proximal single-nucleotide polymorphisms (SNPs) are widely assumed to be independent but could be correlated. Here we introduce a new method, linkage disequilibrium SNP-pair effect correlation regression (LDSPEC), to estimate the correlation of causal disease effect sizes of derived alleles between proximal SNPs; LDSPEC produced robust estimates in simulations. Analyzing 70 UK Biobank diseases and traits (average N = 305,646), we detected significantly non-zero SNP-pair effect correlations (for example, -0.37 ± 0.09 for low-frequency positive linkage disequilibrium 0-100-bp SNP pairs) that decayed with distance and varied with allele frequency and linkage disequilibrium between SNPs. SNP pairs with shared functions had stronger effect correlations that spanned longer genomic distances. Consequently, SNP heritability estimates were smaller than estimates of the sum of causal effect size variances across SNPs, particularly for certain functional annotations. We recapitulated our findings via forward simulations involving stabilizing selection, implicating the action of linkage masking, whereby haplotypes containing linked SNPs with opposite effects on disease have reduced effects on fitness and escape negative selection.

Polymorphism, Single Nucleotide↗

Two new balancer chromosomes on mouse chromosome 4 to facilitate functional annotation of human chromosome 1p.

To facilitate genetic screens to identify and maintain recessive mutations that map to the short arm of human chromosome 1, we have utilized chromosome engineering to generate two mouse strains that carry large inversions on the distal region of mouse chromosome 4. The inversion intervals are 16 and 22 cM in size together they cover approximately half of chromosome 4. Since recombination between the wild-type and inversion chromosomes does not occur within these inversion intervals, mutant alleles of genes mapping to this region can be identified and maintained. Therefore, these inversion chromosomes work as balancer chromosomes. These inversions have the additional advantage that they are tagged with genes encoding the visible coat color markers tyrosinase and agouti, and therefore the dosage of the inversion chromosome (+/+, Inv/+, Inv/Inv) can be visually recognized. These inversion strains will be extremely useful for mutagenesis screens that focus on functional annotation of human chromosome 1p.

Animals↗

Functional annotation of mammalian genomic DNA sequence by chemical mutagenesis: a fine-structure genetic mutation map of a 1- to 2-cM segment of mouse chromosome 7 corresponding to human chromosome 11p14-p15.

Eleven independent, recessive, N-ethyl-N-nitrosourea-induced mutations that map to a approximately 1- to 2-cM region of mouse chromosome (Chr) 7 homologous to human Chr 11p14-p15 were recovered from a screen of 1,218 gametes. These mutations were initially identified in a hemizygous state opposite a large p-locus deletion and subsequently were mapped to finer genomic intervals by crosses to a panel of smaller p deletions. The 11 mutations also were classified into seven complementation groups by pairwise crosses. Four complementation groups were defined by seven prenatally lethal mutations, including a group (l7R3) comprised of two alleles of obvious differing severity. Two allelic mutations (at the psrt locus) result in a severe seizure and runting syndrome, but one mutation (at the fit2 locus) results in a more benign runting phenotype. This experiment has added seven loci, defined by phenotypes of presumed point mutations, to the genetic map of a small (1-2 cM) region of mouse Chr 7 and will facilitate the task of functional annotation of DNA sequence and transcription maps both in the mouse and the corresponding human 11p14-p15 homology region.

Animals↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Functional annotation prediction: all for one and one for all.

In an era of rapid genome sequencing and high-throughput technology, automatic function prediction for a novel sequence is of utter importance in bioinformatics. While automatic annotation methods based on local alignment searches can be simple and straightforward, they suffer from several drawbacks, including relatively low sensitivity and assignment of incorrect annotations that are not associated with the region of similarity. ProtoNet is a hierarchical organization of the protein sequences in the UniProt database. Although the hierarchy is constructed in an unsupervised automatic manner, it has been shown to be coherent with several biological data sources. We extend the ProtoNet system in order to assign functional annotations automatically. By leveraging on the scaffold of the hierarchical classification, the method is able to overcome some frequent annotation pitfalls.

Algorithms↗

SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.

The explosion in the number of functional genomic datasets generated with tools such as DNA microarrays has created a critical need for resources that facilitate the interpretation of large-scale biological data. SOURCE is a web-based database that brings together information from a broad range of resources, and provides it in manner particularly useful for genome-scale analyses. SOURCE's GeneReports include aliases, chromosomal location, functional descriptions, GeneOntology annotations, gene expression data, and links to external databases. We curate published microarray gene expression datasets and allow users to rapidly identify sets of co-regulated genes across a variety of tissues and a large number of conditions using a simple and intuitive interface. SOURCE provides content both in gene and cDNA clone-centric pages, and thus simplifies analysis of datasets generated using cDNA microarrays. SOURCE is continuously updated and contains the most recent and accurate information available for human, mouse, and rat genes. By allowing dynamic linking to individual gene or clone reports, SOURCE facilitates browsing of large genomic datasets. Finally, SOURCEs batch interface allows rapid extraction of data for thousands of genes or clones at once and thus facilitates statistical analyses such as assessing the enrichment of functional attributes within clusters of genes. SOURCE is available at http://source.stanford.edu.

Animals↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Conserved spatially interacting motifs of protein superfamilies: application to fold recognition and function annotation of genome data.

Limitations in techniques for the elucidation of protein function have led to an increasing gap between the annotated proteins and those encoded in a genome. The functional selection and three-dimensional structural constraints of proteins in nature often relate to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. We identify spatially interacting conserved regions, or motifs, within protein superfamilies that are critical for structure and/or function. A search in sequence databases using these descriptors as additional constraints is an approach to identifying putative additional members of superfamilies. Such constrained searches have been tested against proteins of known structure to demonstrate high percentage specificity (93) with a low error rate of 0.0004. This approach has been compared with other sensitive sequence search methods (e.g., PSI-BLAST, HMMsearch, and IMPALA). It has been extended to analyze the distribution of 11 superfamilies in 93 genomes, including the human genome.

Amino Acid Motifs↗

Global profiling of Shewanella oneidensis MR-1: expression of hypothetical genes and improved functional annotations.

The gamma-proteobacterium Shewanella oneidensis strain MR-1 is a metabolically versatile organism that can reduce a wide range of organic compounds, metal ions, and radionuclides. Similar to most other sequenced organisms, approximately 40% of the predicted ORFs in the S. oneidensis genome were annotated as uncharacterized "hypothetical" genes. We implemented an integrative approach by using experimental and computational analyses to provide more detailed insight into gene function. Global expression profiles were determined for cells after UV irradiation and under aerobic and suboxic growth conditions. Transcriptomic and proteomic analyses confidently identified 538 hypothetical genes as expressed in S. oneidensis cells both as mRNAs and proteins (33% of all predicted hypothetical proteins). Publicly available analysis tools and databases and the expression data were applied to improve the annotation of these genes. The annotation results were scored by using a seven-category schema that ranked both confidence and precision of the functional assignment. We were able to identify homologs for nearly all of these hypothetical proteins (97%), but could confidently assign exact biochemical functions for only 16 proteins (category 1; 3%). Altogether, computational and experimental evidence provided functional assignments or insights for 240 more genes (categories 2-5; 45%). These functional annotations advance our understanding of genes involved in vital cellular processes, including energy conversion, ion transport, secondary metabolism, and signal transduction. We propose that this integrative approach offers a valuable means to undertake the enormous challenge of characterizing the rapidly growing number of hypothetical proteins with each newly sequenced genome.

Gene Expression Profiling↗

Comparison of protein active site structures for functional annotation of proteins and drug design.

Rapid and accurate functional assignment of novel proteins is increasing in importance, given the completion of numerous genome sequencing projects and the vastly expanding list of unannotated proteins. Traditionally, global primary-sequence and structure comparisons have been used to determine putative function. These approaches, however, do not emphasize similarities in active site configurations that are fundamental to a protein's activity and highly conserved relative to the global and more variable structural features. The Comparison of Protein Active Site Structures (CPASS) database and software enable the comparison of experimentally identified ligand-binding sites to infer biological function and aid in drug discovery. The CPASS database comprises the ligand-defined active sites identified in the protein data bank, where the CPASS program compares these ligand-defined active sites to determine sequence and structural similarity without maintaining sequence connectivity. CPASS will compare any set of ligand-defined protein active sites, irrespective of the identity of the bound ligand.

Adenosine Triphosphate↗