Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

GObar: a gene ontology based analysis and visualization tool for gene sets.

BACKGROUND: Microarray experiments, as well as other genomic analyses, often result in large gene sets containing up to several hundred genes. The biological significance of such sets of genes is, usually, not readily apparent. Identification of the functions of the genes in the set can help highlight features of interest. The Gene Ontology Consortium 1 has annotated genes in several model organisms using a controlled vocabulary of terms and placed the terms on a Gene Ontology (GO), which comprises three disjoint hierarchies for Molecular functions, Biological processes and Cellular locations. The annotations can be used to identify functions that are enriched in the set, but this analysis can be misleading since the underlying distribution of genes among various functions is not uniform. For example, a large number of genes in a set might be kinases just because the genome contains many kinases. RESULTS: We use the Gene Ontology hierarchy and the annotations to pick significant functions and pathways by comparing the distribution of functions in a given gene list against the distribution of all the genes in the genome, using the hypergeometric distribution to assign probabilities. GObar is a web-based visualizer that implements this algorithm. The public website for GObar 2 can analyse gene lists from the yeast (S. cervisiae), fly (D. Melanogaster), mouse (M. musculus) and human (H. sapiens) genomes. It also allows visualization of the GO tree, as well as placement of a single gene on the GO hierarchy. We analyse a gene list from a genomic study of pre-mRNA splicing to demonstrate the utility of GObar. CONCLUSION: GObar is freely available as a web-based tool at http://katahdin.cshl.org:9331/GO2 and can help analyze and visualize gene lists from genomic analyses.

Algorithms↗

Predicting protein subcellular localization: past, present, and future.

Functional characterization of every single protein is a major challenge of the post-genomic era. The large-scale analysis of a cell's proteins, proteomics, seeks to provide these proteins with reliable annotations regarding their interaction partners and functions in the cellular machinery. An important step on this way is to determine the subcellular localization of each protein. Eukaryotic cells are divided into subcellular compartments, or organelles. Transport across the membrane into the organelles is a highly regulated and complex cellular process. Predicting the subcellular localization by computational means has been an area of vivid activity during recent years. The publicly available prediction methods differ mainly in four aspects: the underlying biological motivation, the computational method used, localization coverage, and reliability, which are of importance to the user. This review provides a short description of the main events in the protein sorting process and an overview of the most commonly used methods in this field.

Computational Biology↗

Genome sequence completed of Alcanivorax borkumensis, a hydrocarbon-degrading bacterium that plays a global role in oil removal from marine systems.

In this paper, we provide background to the genome sequencing project of Alcanivorax borkumensis, which is a marine bacterium that uses exclusively petroleum oil hydrocarbons as sources of carbon and energy (therefore designated "hydrocarbonoclastic"). It is found in low numbers in all oceans of the world and in high numbers in oil-contaminated waters. Its ubiquity and unusual physiology suggest it is globally important in the removal of hydrocarbons from polluted marine systems. A functional genomics analysis of Alcanivorax borkumensis strain SK2 was recently initiated, and its genome sequence has just been completed. Annotation of the genome, metabolome modelling, and functional genomics, will soon reveal important insights into the genomic basis of the properties and physiology of this fascinating and globally important bacterium.

Biodegradation, Environmental↗

Identification of divergent functions in homologous proteins by induction over conserved modules.

Homologous proteins do not necessarily exhibit identical biochemical function. Despite this fact, local or global sequence similarity is widely used as an indication of functional identity. Of the 1327 Enzyme Commission defined functional classes with more than one annotated example in the sequence databases, similarity scores alone are inadequate in 251 (19%) of the cases. We test the hypothesis that conserved domains, as defined in the ProDom database, can be used to discriminate between alternative functions for homologous proteins in these cases. Using machine learning methods, we were able to induce correct discriminators for more than half of these 251 challenging functional classes. These results show that the combination of modular representations of proteins with sequence similarity improves the ability to infer function from sequence over similarity scores alone.

Alcohol Dehydrogenase↗

ProRule: a new database containing functional and structural information on PROSITE profiles.

MOTIVATION: Increase the discriminatory power of PROSITE profiles to facilitate function determination and provide biologically relevant information about domains detected by profiles for the annotation of proteins. SUMMARY: We have created a new database, ProRule, which contains additional information about PROSITE profiles. ProRule contains notably the position of structurally and/or functionally critical amino acids, as well as the condition they must fulfill to play their biological role. These supplementary data should help function determination and annotation of the UniProt Swiss-Prot knowledgebase. ProRule also contains information about the domain detected by the profile in the Swiss-Prot line format. Hence, ProRule can be used to make Swiss-Prot annotation more homogeneous and consistent. The format of ProRule can be extended to provide information about combination of domains. AVAILABILITY: ProRule can be accessed through ScanProsite at http://www.expasy.org/tools/scanprosite. A file containing the rules will be made available under the PROSITE copyright conditions on our ftp site (ftp://www.expasy.org/databases/prosite/) by the next PROSITE release.

Amino Acid Sequence↗

MultiFun, a multifunctional classification scheme for Escherichia coli K-12 gene products.

An enriched classification system for cellular functions of gene products of Escherichia coli K-12 was developed based on the initial classification by Riley. In the new classification scheme, MultiFun, cellular functions are divided into 10 major categories: Metabolism, Information Transfer, Regulation, Transport, Cell Processes, Cell Structure, Location, Extra-chromosomal Origin, DNA Site, and Cryptic Gene. These major categories are further sub-divided into a hierarchical scheme. Two thousand nine hundred twenty-two gene products of E. coli K-12 were assigned to one or more functions depending on the role they play in the cell. Functional assignments were made to 66% of E. coli gene products, ranging from 1 to 16 assignments per gene product. The expansion of cellular function categories and the assignment to more than one category (multifunction) provides a more complete description of the gene products and their roles and hence better reflects the functional complexity of organisms. We believe this classification system will be useful in the field of genome analysis, both for annotation purposes and for comparative studies. The functional classification scheme and the cellular function assignments made to E. coli gene products can be accessed from the web at the databases GenProtEC (http://genprotec.mbl.edu) and EcoCyc (http://www.ecocyc.org).

Bacterial Proteins↗

GOChase: correcting errors from Gene Ontology-based annotations for gene products.

SUMMARY: The Gene Ontology (GO) is a controlled biological vocabulary that provides three structured networks of terms to describe biological processes, cellular components and molecular functions. Many databases of gene products are annotated using the GO vocabularies. We found that some GO-updating operations are not easily traceable by the current biological databases and GO browsers. Consequently, numerous annotation errors arise and are propagated throughout biological databases and GO-based high-level analyses. GOChase is a set of web-based utilities to detect and correct the errors in GO-based annotations.

Database Management Systems↗

Glycosyltransferases in SWISS-PROT.

SWISS-PROT is a curated protein sequence database with a high level of annotation (such as description of the function of a protein, its domain structure, post-translational modification, variants, etc), a minimal level of redundancy and a high level of integration with other databases. An ongoing project is to maintain the glycosyltransferase family of enzymes with comprehensive annotation and documentation in the SWISS-PROT database and to represent the most recent research developments.

Amino Acid Sequence↗

VDJ-Insights: simplifying the annotation of genomic immunoglobulin and T cell receptor regions.

MOTIVATION: Accurate annotation of germline immunoglobulin (IG) and T cell receptor (TCR) loci is critical for understanding adaptive immunity. RESULTS: VDJ-Insights provides a user-friendly software package for characterizing these complex immune regions. In addition, it assesses gene segment functionality, identifies recombination signal sequences, and annotates complementarity-determining regions 1 and 2. VDJ-Insights achieved over 99% concordance with curated annotations from multiple species, outperforming existing annotation tools. When applied to 95 haplotypes from the Human Pangenome Reference Consortium, VDJ-Insights identified 652 and 275 novel IG and TCR alleles, respectively, highlighting its scalability for large immunogenetic studies. AVAILABILITY AND IMPLEMENTATION: Datasets and software package are available in the VDJ-insights repository, https://github.com/BPRC-Bioinfo and https://doi.org/10.5281/zenodo.17588835. Additional intermediate datasets used and analyzed during the current study are available from the corresponding authors upon reasonable request.

Software↗

Genomewide function conservation and phylogeny in the Herpesviridae.

The Herpesviridae are a large group of well-characterized double-stranded DNA viruses for which many complete genome sequences have been determined. We have extracted protein sequences from all predicted open reading frames of 19 herpesvirus genomes. Sequence comparison and protein sequence clustering methods have been used to construct herpesvirus protein homologous families. This resulted in 1692 proteins being clustered into 243 multiprotein families and 196 singleton proteins. Predicted functions were assigned to each homologous family based on genome annotation and published data and each family classified into seven broad functional groups. Phylogenetic profiles were constructed for each herpesvirus from the homologous protein families and used to determine conserved functions and genomewide phylogenetic trees. These trees agreed with molecular-sequence-derived trees and allowed greater insight into the phylogeny of ungulate and murine gammaherpesviruses.

Animals↗

Sputnik: a database platform for comparative plant genomics.

Two million plant ESTs, from 20 different plant species, and totalling more than one 1000 Mbp of DNA sequence, represents a formidable transcriptomic resource. Sputnik uses the potential of this sequence resource to fill some of the information gap in the un-sequenced plant genomes and to serve as the foundation for in silicio comparative plant genomics. The complexity of the individual EST collections has been reduced using optimised EST clustering techniques. Annotation of cluster sequences is performed by exploiting and transferring information from the comprehensive knowledgebase already produced for the completed model plant genome (Arabidopsis thaliana) and by performing additional state of-the-art sequence analyses relevant to today's plant biologist. Functional predictions, comparative analyses and associative annotations for 500 000 plant EST derived peptides make Sputnik (http://mips.gsf.de/proj/sputnik/) a valid platform for contemporary plant genomics.

Databases, Nucleic Acid↗

The past, present and future of genome-wide re-annotation.

Annotation, the process by which structural or functional information is inferred for genes or proteins, is crucial for obtaining value from genome sequences. We define the process of annotating a previously annotated genome sequence as 're-annotation', and examine the strengths and weaknesses of current manual and automatic genome-wide re-annotation approaches.

Computational Biology↗

A semantic analysis of the annotations of the human genome.

The correct interpretation of any biological experiment depends in an essential way on the accuracy and consistency of the existing annotation databases. Such databases are ubiquitous and used by all life scientists in most experiments. However, it is well known that such databases are incomplete and many annotations may also be incorrect. In this paper we describe a technique that can be used to analyze the semantic content of such annotation databases. Our approach is able to extract implicit semantic relationships between genes and functions. This ability allows us to discover novel functions for known genes. This approach is able to identify missing and inaccurate annotations in existing annotation databases, and thus help improve their accuracy. We used our technique to analyze the current annotations of the human genome. From this body of annotations, we were able to predict 212 additional gene-function assignments. A subsequent literature search found that 138 of these gene-functions assignments are supported by existing peer-reviewed papers. An additional 23 assignments have been confirmed in the meantime by the addition of the respective annotations in later releases of the Gene Ontology database. Overall, the 161 confirmed assignments represent 75.95% of the proposed gene-function assignments. Only one of our predictions (0.4%) was contradicted by the existing literature. We could not find any relevant articles for 50 of our predictions (23.58%). The method is independent of the organism and can be used to analyze and improve the quality of the data of any public or private annotation database.

Chromosome Mapping↗

An EST-based approach for identifying genes expressed in the intestine and gills of pre-smolt Atlantic salmon (Salmo salar).

BACKGROUND: The Atlantic salmon is an important aquaculture species and a very interesting species biologically, since it spawns in fresh water and develops through several stages before becoming a smolt, the stage at which it migrates to the sea to feed. The dramatic change of habitat requires physiological, morphological and behavioural changes to prepare the salmon for its new environment. These changes are called the parr-smolt transformation or smoltification, and pre-adapt the salmon for survival and growth in the marine environment. The development of hypo-osmotic regulatory ability plays an important part in facilitating the transition from rivers to the sea. The physiological mechanisms behind the developmental changes are largely unknown. An understanding of the transformation process will be vital to the future of the aquaculture industry. A knowledge of which genes are expressed prior to the smoltification process is an important basis for further studies. RESULTS: In all, 2974 unique sequences, consisting of 779 contigs and 2195 singlets, were generated for Atlantic salmon from two cDNA libraries constructed from the gills and the intestine, accession numbers [Genbank: CK877169-CK879929, CK884015-CK886537 and CN181112-CN181464]. Nearly 50% of the sequences were assigned putative functions because they showed similarity to known genes, mostly from other species, in one or more of the databases used. The Swiss-Prot database returned significant hits for 1005 sequences. These could be assigned predicted gene products, and 967 were annotated using Gene Ontology (GO) terms for molecular function, biological process and/or cellular component, employing an annotation transfer procedure. CONCLUSION: This paper describes the construction of two cDNA libraries from pre-smolt Atlantic salmon (Salmo salar) and the subsequent EST sequencing, clustering and assigning of putative function to 1005 genes expressed in the gills and/or intestine.

Animals↗

Globin gene server: a prototype E-mail database server featuring extensive multiple alignments and data compilation for electronic genetic analysis.

The sequence of virtually the entire cluster of beta-like globin genes has been determined from several mammals, and many regulatory regions have been analyzed by mutagenesis, functional assays, and nuclear protein binding studies. This very large amount of sequence and functional data needs to be compiled in a readily accessible and usable manner to optimize data analysis, hypothesis testing, and model building. We report a Globin Gene Server that will provide this service in a constantly updated manner when fully implemented. The Server has two principal functions. The first (currently available) provides an annotated multiple alignment of the DNA sequences throughout the gene cluster from representatives of all species analyzed. The second compiles data on functional and protein binding assays throughout the gene cluster. A prototype of this compilation using the aligned 5' flanking region of beta-globin genes from five species shows examples of (1) well-conserved regions that have demonstrated functions, including cases in which the functional data are in apparent conflict, (2) proposed functional regions that are not well conserved, and (3) conserved regions with no currently assigned function. Such an electronic genetic analysis leads to many readily testable hypotheses that were not immediately apparent without the multiple alignment and compilation. The Server is accessible via E-mail on computer networks, and printed results can be obtained by request to the authors. This prototype will be a helpful guide for developing similar tools for many genomic loci.

Animals↗

Development of a functional genomics platform for Sinorhizobium meliloti: construction of an ORFeome.

The nitrogen-fixing, symbiotic bacterium Sinorhizobium meliloti reduces molecular dinitrogen to ammonia in a specific symbiotic context, supporting the nitrogen requirements of various forage legumes, including alfalfa. Determining the DNA sequence of the S. meliloti genome was an important step in plant-microbe interaction research, adding to the considerable information already available about this bacterium by suggesting possible functions for many of the >6,200 annotated open reading frames (ORFs). However, the predictive power of bioinformatic analysis is limited, and putting the role of these genes into a biological context will require more definitive functional approaches. We present here a strategy for genetic analysis of S. meliloti on a genomic scale and report the successful implementation of the first step of this strategy by constructing a set of plasmids representing 100% of the 6,317 annotated ORFs cloned into a mobilizable plasmid by using efficient PCR and recombination protocols. By using integrase recombination to insert these ORFs into other plasmids in vitro or in vivo (B. L. House et al., Appl. Environ. Microbiol. 70:2806-2815, 2004), this ORFeome can be used to generate various specialized genetic materials for functional analysis of S. meliloti, such as operon fusions, mutants, and protein expression plasmids. The strategy can be generalized to many other genome projects, and the S. meliloti clones should be useful for investigators wanting an accessible source of cloned genes encoding specific enzymes.

Bacterial Proteins↗

Integration of GO annotations in Correspondence Analysis: facilitating the interpretation of microarray data.

MOTIVATION: The functional interpretation of microarray datasets still represents a time-consuming and challenging task. Up to now functional categories that are relevant for one or more experimental context(s) have been commonly extracted from a set of regulated genes and presented in long lists. RESULTS: To facilitate interpretation, we integrated Gene Ontology (GO) annotations into Correspondence Analysis to display genes, experimental conditions and gene-annotations in a single plot. The position of the annotations in these plots can be directly used for the functional interpretation of clusters of genes or experimental conditions without the need for comparing long lists of annotations. Correspondence Analysis is not limited in the number of experimental conditions that can be compared simultaneously, allowing an easy identification of characterizing annotations even in complex experimental settings. Due to the rapidly increasing amount of annotation data available, we apply an annotation filter. Hereby the number of displayed annotations can be significantly reduced to a set of descriptive ones, further enhancing the interpretability of the plot. We validated the method on transcription data from Saccharomyces cerevisiae and human pancreatic adenocarcinomas. AVAILABILITY: The M-CHiPS software is accessible for collaborators at http://www.mchips.org

Algorithms↗

Complementing computationally predicted regulatory sites in Tractor_DB using a pattern matching approach.

Prokaryotic genomes annotation has focused on genes location and function. The lack of regulatory information has limited the knowledge on cellular transcriptional regulatory networks. However, as more phylogenetically close genomes are sequenced and annotated, the implementation of phylogenetic footprinting strategies for the recognition of regulators and their regulons becomes more important. In this paper we describe a comparative genomics approach to the prediction of new gamma-proteobacterial regulon members. We take advantage of the phylogenetic proximity of Escherichia coli and other 16 organisms of this subdivision and the intensive search of the space sequence provided by a pattern-matching strategy. Using this approach we complement predictions of regulatory sites made using statistical models currently stored in Tractor_DB, and increase the number of transcriptional regulators with predicted binding sites up to 86. All these computational predictions may be reached at Tractor_DB (www.bioinfo.cu/Tractor_DB, www.tractor.lncc.br, www.ccg.unam.mx/Computational_Genomics/tractorDB/). We also take a first step in this paper towards the assessment of the conservation of the architecture of the regulatory network in the gamma-proteobacteria through evaluating the conservation of the overall connectivity of the network.

Base Sequence↗