Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Functional annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Automatic annotation of protein function.

The annotation of protein function at genomic scale is essential for day-to-day work in biology and for any systematic approach to the modeling of biological systems. Currently, functional annotation is essentially based on the expansion of the relatively small number of experimentally determined functions to large collections of proteins. The task of systematic annotation faces formidable practical problems related to the accuracy of the input experimental information, the reliability of current systems for transferring information between related sequences, and the reproducibility of the links between database information and the original experiments reported in publications. These technical difficulties merely lie on the surface of the deeper problem of the evolution of protein function in the context of protein sequences and structures. Given the mixture of technical and scientific challenges, it is not surprising that errors are introduced, and expanded, in database annotations. In this situation, a more realistic option is the development of a reliability index for database annotations, instead of depending exclusively on efforts to correct databases. Several groups have attempted to compare the database annotations of similar proteins, which constitutes the first steps toward the calibration of the relationship between sequence and annotation space.

Artificial Intelligence↗

An integrated analysis and database system for full-length cDNA.

Annotation and database system of full-length cDNA sequences was developed. As the components of the system, ORF annotation system, functional annotation system based on database search results, mapping annotation system, and integrated retrieval and display system were developed. In the ORF annotation system integrated analyses using conventional tools are performed and useful retrieval interface using motif list are introduced. In the functional annotation system based on database search results, a new method that characterizes a given unknown cDNA was developed by using a profile of similarity level over words appearing in sequence database entries. In the mapping annotation system, we linked by similarity searches full-length cDNA sequences with database DNA sequences that are already mapped on chromosomes. By using these links, full-length cDNAs can be retrieved by the retrieval condition of physical mapping information. Genetic disease information mapped on the physical mapping site can also be displayed by this system. Furthermore, we constructed an integrated database system for these analyzed data, and thus enabled annotation and selection of full-length cDNAs from points of both gene function and mapping information.

Chromosome Mapping↗

Genome-wide association study reveals candidate genes associated with body weight and wool traits in Ordos fine-wool sheep.

BACKGROUND: The Ordos fine-wool sheep is a high-quality fine-wool breed in China, renowned for its excellent wool quality, meat production, and adaptability to the arid and semi-arid regions of Inner Mongolia. Body weight and wool traits are important economic characteristics in sheep breeding. This study aimed to identify genetic loci associated with body weight (BW), wool length (WL), and wool fineness (WF) in Ordos fine-wool sheep. METHODS: A genome-wide association study (GWAS) was conducted in 388 Ordos fine-wool sheep genotyped using the GenoBaits® Ovine 40K SNP panel. Single nucleotide polymorphisms (SNPs) associated with BW, WL, and WF were identified, and candidate genes located near the SNPs reaching the suggestive threshold were subjected to functional annotation and enrichment analysis. RESULTS: A total of 22 SNPs were identified as potentially associated with BW, WL, and WF traits, corresponding to 27 annotated genes. Functional annotation highlighted six potential candidate genes, including LAMA2, ARHGAP18, IGFBP2, IGFBP5, CA10, and AXIN1, which may play important roles in regulating body weight and wool growth in sheep. CONCLUSIONS: The identified genes provide valuable candidate loci for BW, WL, and WF traits in Ordos fine-wool sheep. The results of this study provide preliminary references for further exploration of the genetic mechanisms of wool traits in Ordos fine-wool sheep and the development of molecular breeding markers.

GWAS↗

SNP Function Portal: a web database for exploring the function implication of SNP alleles.

MOTIVATION: Finding the potential functional significance of SNPs is a major bottleneck in understanding genome-wide SNP scanning results, as the related functional data are distributed across many different databases. The SNP Function Portal is designed to be a clearing house for all public domain SNP functional annotation data, as well as in-house functional annotations derived from different data sources. It currently contains SNP functional annotations in six major categories including genomic elements, transcription regulation, protein function, pathway, disease and population genetics. Besides extensive SNP functional annotations, the SNP Function Portal includes a powerful search engine that accepts different types of genetic markers as input and identifies all genetically related SNPs based on the HapMap Phase II data as well as the relationship of different markers to known genes. As a result, our system allows users to identify the potential biological impact of genetic markers and complex relationships among genetic markers and genes, and it greatly facilitates knowledge discovery in genome-wide SNP scanning experiments. AVAILABILITY: http://brainarray.mbni.med.umich.edu/Brainarray/Database/SearchSNP/snpfunc.aspx.

Alleles↗

Functional genome annotation through phylogenomic mapping.

Accurate determination of functional interactions among proteins at the genome level remains a challenge for genomic research. Here we introduce a genome-scale approach to functional protein annotation--phylogenomic mapping--that requires only sequence data, can be applied equally well to both finished and unfinished genomes, and can be extended beyond single genomes to annotate multiple genomes simultaneously. We have developed and applied it to more than 200 sequenced bacterial genomes. Proteins with similar evolutionary histories were grouped together, placed on a three dimensional map and visualized as a topographical landscape. The resulting phylogenomic maps display thousands of proteins clustered in mountains on the basis of coinheritance, a strong indicator of shared function. In addition to systematic computational validation, we have experimentally confirmed the ability of phylogenomic maps to predict both mutant phenotype and gene function in the delta proteobacterium Myxococcus xanthus.

Bacterial Proteins↗

Quantifying the relationship between co-expression, co-regulation and gene function.

BACKGROUND: It is thought that genes with similar patterns of mRNA expression and genes with similar functions are likely to be regulated via the same mechanisms. It has been difficult to quantitatively test these hypotheses on a large scale because there has been no general way of determining whether genes share a common regulatory mechanism. Here we use data from a recent genome wide binding analysis in combination with mRNA expression data and existing functional annotations to quantify the likelihood that genes with varying degrees of similarity in mRNA expression profile or function will be bound by a common transcription factor. RESULTS: Genes with strongly correlated mRNA expression profiles are more likely to have their promoter regions bound by a common transcription factor. This effect is present only at relatively high levels of expression similarity. In order for two genes to have a greater than 50% chance of sharing a common transcription factor binder, the correlation between their expression profiles (across the 611 microarrays used in our study) must be greater than 0.84. Genes with similar functional annotations are also more likely to be bound by a common transcription factor. Combining mRNA expression data with functional annotation results in a better predictive model than using either data source alone. CONCLUSIONS: We demonstrate how mRNA expression data and functional annotations can be used together to estimate the probability that genes share a common regulatory mechanism. Existing microarray data and known functional annotations are sufficient to identify only a relatively small percentage of co-regulated genes.

Gene Expression Profiling↗

Transcript annotation in FANTOM3: mouse gene catalog based on physical cDNAs.

The international FANTOM consortium aims to produce a comprehensive picture of the mammalian transcriptome, based upon an extensive cDNA collection and functional annotation of full-length enriched cDNAs. The previous dataset, FANTOM2, comprised 60,770 full-length enriched cDNAs. Functional annotation revealed that this cDNA dataset contained only about half of the estimated number of mouse protein-coding genes, indicating that a number of cDNAs still remained to be collected and identified. To pursue the complete gene catalog that covers all predicted mouse genes, cloning and sequencing of full-length enriched cDNAs has been continued since FANTOM2. In FANTOM3, 42,031 newly isolated cDNAs were subjected to functional annotation, and the annotation of 4,347 FANTOM2 cDNAs was updated. To accomplish accurate functional annotation, we improved our automated annotation pipeline by introducing new coding sequence prediction programs and developed a Web-based annotation interface for simplifying the annotation procedures to reduce manual annotation errors. Automated coding sequence and function prediction was followed with manual curation and review by expert curators. A total of 102,801 full-length enriched mouse cDNAs were annotated. Out of 102,801 transcripts, 56,722 were functionally annotated as protein coding (including partial or truncated transcripts), providing to our knowledge the greatest current coverage of the mouse proteome by full-length cDNAs. The total number of distinct non-protein-coding transcripts increased to 34,030. The FANTOM3 annotation system, consisting of automated computational prediction, manual curation, and final expert curation, facilitated the comprehensive characterization of the mouse transcriptome, and could be applied to the transcriptomes of other species.

Animals↗

ELISA: a unified, multidimensional view of the protein domain universe.

ELISA (http://romi.bu.edu/elisa/) is a database that was designed for flexibility in defining interesting queries about protein domain evolution. We have defined and included both the inherent characteristics of the domains such as structure and function and comparisons of these characteristics between domains. Thus, the database is useful in defining structural and functional links between related protein domains and by extension sequences that encode them. In this database we introduce and employ a novel method of functional annotation and comparison. For each protein domain we create a probabilistic functional annotation tree using GO. We have designed an algorithm that accurately compares these trees and thus provides a measure of "functional distance" between two protein domains. Along with functional annotation, we have also included structural comparison between protein domains and best sequence comparisons to all known genomes. The latter enables researchers to dynamically do searches for domains sharing similar phylogenetic profiles. This combination of data and tools enables the researcher to design complex queries to carry out research in the areas of protein domain evolution, structure prediction and functional annotation of novel sequences.

Amino Acid Sequence↗

Cause and effect considerations in diagnostic pathology and pathology phenotyping of genetically engineered mice (GEM).

Over the next several decades, biology is embarking on its most ambitious project yet: to annotate the human genome functionally, prioritizing and focusing on those genes relevant to development and disease. Model systems are fundamental prerequisites for this task, and genetically engineered mice (GEM) are by far the most accessible mammalian system because of their anatomical, physiological, and genetic similarity to humans. The scientific utility of GEM has become commonplace since the technology to produce them was established in the early 1980s. Conceptually, however, an efficiently coordinated high-throughput approach that permits correlation between newly discovered genes, functional properties of their protein products, and biological relevance of these products as drug targets has yet to be established. The discipline of veterinary anatomical pathology (hereafter referred to as pathology) is not immune to this requirement for evolution and adaptation, and to address relationships and tissue consequences between tens of thousands of genes and their cognate proteins, novel interdisciplinary technologies and approaches must emerge. Although many of the techniques of pathology are well established, in the context of pathology's contribution to functional annotation of the genome, several conceptually important and unresolved issues remain to be addressed. While an ever-increasing arsenal of genetic and molecular tool-sets are available to evaluate and understand the function of genes and their pathophysiological mechanisms, pathology will continue to play an essential role in confirming cause and effect relationships of gene function in development and disease. This role will continue to be dependent on keen observation, a systematic but disciplined approach, expert knowledge of strain-dependent anatomical differences and incidental lesions, and relevant tissue-based evidence. Miniaturization and high-throughput adaptation of these methods must also continue so that they can complement parallel phenotyping efforts, provide pathology-based data in pace with concurrent phenotyping efforts, and continue to find new utility in the collective effort of functional annotation.

Animals↗

Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.

We have developed a prototype for the automatic annotation of functional characteristics in protein families. The system is able to extract biological information directly from scientific literature in the form of MEDLINE abstracts. The criterion for selecting relevant keywords is the difference between their frequency in the abstracts associated with the protein family under study and its frequency in other unrelated protein families. The concept of functional information associated to protein families is the key feature of our system and gathers evolutionary information into the problem of functional annotation of biological sequences. The system has been tested in two different scenarios: first, a large set of protein families with a small number of abstract per family and second, selected protein families with large number of abstracts attached to each one. In both cases the performances are compared with annotations provided by human experts showing a clear relation between the amount of information provided to the system and the quality of the annotations. The automatic annotations are in many cases of similar quality to the ones contained in current data bases. The possibilities and difficulties to be encountered during the development of a full system for automatic annotation are discussed.

Abstracting and Indexing↗

Modeling the percolation of annotation errors in a database of protein sequences.

Public sequence databases contain information on the sequence, structure and function of proteins. Genome sequencing projects have led to a rapid increase in protein sequence information, but reliable, experimentally verified, information on protein function lags a long way behind. To address this deficit, functional annotation in protein databases is often inferred by sequence similarity to homologous, annotated proteins, with the attendant possibility of error. Now, the functional annotation in these homologous proteins may itself have been acquired through sequence similarity to yet other proteins, and it is generally not possible to determine how the functional annotation of any given protein has been acquired. Thus the possibility of chains of misannotation arises, a process we term 'error percolation'. With some simple assumptions, we develop a dynamical probabilistic model for these misannotation chains. By exploring the consequences of the model for annotation quality it is evident that this iterative approach leads to a systematic deterioration of database quality.

Animals↗

FSSA: a novel method for identifying functional signatures from structural alignments.

MOTIVATION: It is commonly believed that sequence determines structure, which in turn determines function. However, the presence of many proteins with the same structural fold but different functions suggests that global structure and function do not always correlate well. RESULTS: We propose a method for accurate functional annotation, based on identification of functional signatures from structural alignments (FSSA) using the Structural Classification of Proteins (SCOP) database. The FSSA method is superior at function discrimination and classification compared with several methods that directly inherit functional annotation information from homology inference, such as Smith-Waterman, PSI-BLAST, hidden Markov models and structure comparison methods, for a large number of structural fold families. Our results indicate that the contributions of amino acid residue types and positions to structure and function are largely separable for proteins in multi-functional fold families.

Algorithms↗

DIG--a system for gene annotation and functional discovery.

SUMMARY: We describe a database and information discovery system named DIG (Duke Integrated Genomics) designed to facilitate the process of gene annotation and the discovery of functional context. The DIG system collects and organizes gene annotation and functional information, and includes tools that support an understanding of genes in a functional context by providing a framework for integrating and visualizing gene expression, protein interaction and literature-based interaction networks.

Chromosome Mapping↗

A novel deep learning-driven framework for improving lncRNA comprehensive annotation with LncADeep 2.0.

MOTIVATION: Long non-coding RNAs (lncRNAs) have emerged as crucial players in diverse physiological and pathological processes, yet the biological mechanisms of the vast majority of lncRNAs remain elusive. To fill this gap, it is necessary to improve the accuracy of lncRNA identification and functional annotation. RESULTS: Here, we introduce LncADeep 2.0, an integrated deep learning framework designed to meet these needs. In the identification module, LncADeep 2.0 incorporated novel peptide features along with sequence and structural information, demonstrating superior performance over our previous LncADeep and other existing tools on both annotated transcripts from GENCODE and RNA-seq data. For functional annotation, LncADeep 2.0 leveraged lncRNA-centric interaction networks and gene ontology terms through the transfer learning strategy to achieve robust annotation performance with limited functional data. Compared to LncADeep, LncADeep 2.0 could accurately elucidate the general functions of given lncRNA sequences, predict tissue- or cell-type-specific functions from bulk and single-cell RNA-seq data, and establish connections between tumor-associated lncRNAs and genomic markers. Overall, LncADeep 2.0 stands out as an efficient and reliable tool for lncRNA identification and functional annotation across a wide spectrum of biological processes. AVAILABILITY AND IMPLEMENTATION: LncADeep 2.0 is available for use at https://github.com/Jefferson-Chou/LncADeep2 and https://doi.org/10.5281/zenodo.17164767.

RNA, Long Noncoding↗

AgBase: a functional genomics resource for agriculture.

BACKGROUND: Many agricultural species and their pathogens have sequenced genomes and more are in progress. Agricultural species provide food, fiber, xenotransplant tissues, biopharmaceuticals and biomedical models. Moreover, many agricultural microorganisms are human zoonoses. However, systems biology from functional genomics data is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation and agricultural research communities are smaller with limited funding compared to many model organism communities. DESCRIPTION: To facilitate systems biology in these traditionally agricultural species we have established "AgBase", a curated, web-accessible, public resource http://www.agbase.msstate.edu for structural and functional annotation of agricultural genomes. The AgBase database includes a suite of computational tools to use GO annotations. We use standardized nomenclature following the Human Genome Organization Gene Nomenclature guidelines and are currently functionally annotating chicken, cow and sheep gene products using the Gene Ontology (GO). The computational tools we have developed accept and batch process data derived from different public databases (with different accession codes), return all existing GO annotations, provide a list of products without GO annotation, identify potential orthologs, model functional genomics data using GO and assist proteomics analysis of ESTs and EST assemblies. Our journal database helps prevent redundant manual GO curation. We encourage and publicly acknowledge GO annotations from researchers and provide a service for researchers interested in GO and analysis of functional genomics data. CONCLUSION: The AgBase database is the first database dedicated to functional genomics and systems biology analysis for agriculturally important species and their pathogens. We use experimental data to improve structural annotation of genomes and to functionally characterize gene products. AgBase is also directly relevant for researchers in fields as diverse as agricultural production, cancer biology, biopharmaceuticals, human health and evolutionary biology. Moreover, the experimental methods and bioinformatics tools we provide are widely applicable to many other species including model organisms.

Agriculture↗

FIGENIX: intelligent automation of genomic annotation: expertise integration in a new software platform.

BACKGROUND: Two of the main objectives of the genomic and post-genomic era are to structurally and functionally annotate genomes which consists of detecting genes' position and structure, and inferring their function (as well as of other features of genomes). Structural and functional annotation both require the complex chaining of numerous different software, algorithms and methods under the supervision of a biologist. The automation of these pipelines is necessary to manage huge amounts of data released by sequencing projects. Several pipelines already automate some of these complex chaining but still necessitate an important contribution of biologists for supervising and controlling the results at various steps. RESULTS: Here we propose an innovative automated platform, FIGENIX, which includes an expert system capable to substitute to human expertise at several key steps. FIGENIX currently automates complex pipelines of structural and functional annotation under the supervision of the expert system (which allows for example to make key decisions, check intermediate results or refine the dataset). The quality of the results produced by FIGENIX is comparable to those obtained by expert biologists with a drastic gain in terms of time costs and avoidance of errors due to the human manipulation of data. CONCLUSION: The core engine and expert system of the FIGENIX platform currently handle complex annotation processes of broad interest for the genomic community. They could be easily adapted to new, or more specialized pipelines, such as for example the annotation of miRNAs, the classification of complex multigenic families, annotation of regulatory elements and other genomic features of interest.

Automation↗

The wheat (Triticum aestivum L.) leaf proteome.

The wheat leaf proteome was mapped and partially characterized to function as a comparative template for future wheat research. In total, 404 proteins were visualized, and 277 of these were selected for analysis based on reproducibility and relative quantity. Using a combination of protein and expressed sequence tag database searching, 142 proteins were putatively identified with an identification success rate of 51%. The identified proteins were grouped according to their functional annotations with the majority (40%) being involved in energy production, primary, or secondary metabolism. Only 8% of the protein identifications lacked ascertainable functional annotation. The 51% ratio of successful identification and the 8% unclear functional annotation rate are major improvements over most previous plant proteomic studies. This clearly indicates the advancement of the plant protein and nucleic acid sequence and annotation data available in the databases, and shows the enhanced feasibility of future wheat leaf proteome research.

Computational Biology↗

Gene fusions and gene duplications: relevance to genomic annotation and functional analysis.

BACKGROUND: Escherichia coli a model organism provides information for annotation of other genomes. Our analysis of its genome has shown that proteins encoded by fused genes need special attention. Such composite (multimodular) proteins consist of two or more components (modules) encoding distinct functions. Multimodular proteins have been found to complicate both annotation and generation of sequence similar groups. Previous work overstated the number of multimodular proteins in E. coli. This work corrects the identification of modules by including sequence information from proteins in 50 sequenced microbial genomes. RESULTS: Multimodular E. coli K-12 proteins were identified from sequence similarities between their component modules and non-fused proteins in 50 genomes and from the literature. We found 109 multimodular proteins in E. coli containing either two or three modules. Most modules had standalone sequence relatives in other genomes. The separated modules together with all the single (un-fused) proteins constitute the sum of all unimodular proteins of E. coli. Pairwise sequence relationships among all E. coli unimodular proteins generated 490 sequence similar, paralogous groups. Groups ranged in size from 92 to 2 members and had varying degrees of relatedness among their members. Some E. coli enzyme groups were compared to homologs in other bacterial genomes. CONCLUSION: The deleterious effects of multimodular proteins on annotation and on the formation of groups of paralogs are emphasized. To improve annotation results, all multimodular proteins in an organism should be detected and when known each function should be connected with its location in the sequence of the protein. When transferring functions by sequence similarity, alignment locations must be noted, particularly when alignments cover only part of the sequences, in order to enable transfer of the correct function. Separating multimodular proteins into module units makes it possible to generate protein groups related by both sequence and function, avoiding mixing of unrelated sequences. Organisms differ in sizes of groups of sequence-related proteins. A sample comparison of orthologs to selected E. coli paralogous groups correlates with known physiological and taxonomic relationships between the organisms.

Computational Biology↗