Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

The Lipase Engineering Database: a navigation and analysis tool for protein families.

The Lipase Engineering Database (LED) (http://www.led.uni-stuttgart.de) integrates information on sequence, structure, and function of lipases, esterases, and related proteins. Sequence data on 806 protein entries are assigned to 38 homologous families, which are grouped into 16 superfamilies with no global sequence similarity between each other. For each family, multisequence alignments are provided with functionally relevant residues annotated. Pre-calculated phylogenetic trees allow navigation inside superfamilies. Experimental structures of 45 proteins are superposed and consistently annotated. The LED has been applied to systematically analyze sequence-structure-function relationships of this vast and diverse enzyme class. It is a useful tool to identify functionally relevant residues apart from the active site residues, and to design mutants with desired substrate specificity.

Amino Acid Sequence↗

Only one of the two annotated Lactococcus lactis fabG genes encodes a functional beta-ketoacyl-acyl carrier protein reductase.

The small genome of the Gram-positive bacterium Lactococcus lactis ssp. lactis IL1403 contains two genes that encode proteins annotated as homologues of Escherichia coli beta-hydroxyacyl-acyl carrier protein (ACP) reductase. E. coli fabG encodes beta-ketoacyl-acyl carrier protein (ACP) reductase, the enzyme responsible for the first reductive step of the fatty acid synthetic cycle. Both of the L. lactis genes are adjacent to (and predicted to be cotranscribed with) other genes that encode proteins having homology to known fatty acid synthetic enzymes. Such relationships have often been used to strengthen annotations based on sequence alignments. Annotation in the case of beta-ketoacyl-ACP reductase is particularly problematic because the protein is a member of a vast protein family, the short-chain alcohol dehydrogenase/reductase (SDR) family. The recent isolation of an E. coli fabG mutant strain encoding a conditionally active beta-ketoacyl-ACP reductase allowed physiological and biochemical testing of the putative L. lactishomologues. We report that expression of only one of the two L. lactis proteins (that annotated as FabG1) allows growth of the E. coli fabG strain under nonpermissive conditions and restores in vitro fatty acid synthetic ability to extracts of the mutant strain. Therefore, like E. coli, L. lactis has a single beta-ketoacyl-ACP reductase active with substrates of all fatty acid chain lengths. The second protein (annotated as FabG2), although inactive in fatty acid synthesis both in vivo and in vitro, was highly active in reduction of the model substrate, beta-ketobutyryl-CoA. As expected from work on the E. coli enzyme, the FabG1 beta-ketobutyryl-CoA reductase activity was inhibited by ACP (which blocks access to the active site) whereas the activity of FabG2 was unaffected by the presence of ACP. These results seem to be an example of a gene duplication event followed by divergence of one copy of the gene to encode a protein having a new function.

Alcohol Oxidoreductases↗

Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.

Functional characterizations of thousands of gene products from many species are described in the published literature. These discussions are extremely valuable for characterizing the functions not only of these gene products, but also of their homologs in other organisms. The Gene Ontology (GO) is an effort to create a controlled terminology for labeling gene functions in a more precise, reliable, computer-readable manner. Currently, the best annotations of gene function with the GO are performed by highly trained biologists who read the literature and select appropriate codes. In this study, we explored the possibility that statistical natural language processing techniques can be used to assign GO codes. We compared three document classification methods (maximum entropy modeling, naïve Bayes classification, and nearest-neighbor classification) to the problem of associating a set of GO codes (for biological process) to literature abstracts and thus to the genes associated with the abstracts. We showed that maximum entropy modeling outperforms the other methods and achieves an accuracy of 72% when ascertaining the function discussed within an abstract. The maximum entropy method provides confidence measures that correlate well with performance. We conclude that statistical methods may be used to assign GO codes and may be useful for the difficult task of reassignment as terminology standards evolve over time.

Algorithms↗

Social behavior of the yeast protein-protein interaction network.

Protein-protein interaction networks are useful in contextual annotation of protein function and in general to achieve a system-level understanding of cellular behavior. This work reports on the social behavior of the yeast protein-protein interaction network and concludes that it is non-random. This work, while providing an analysis of organization of genes into functional societies, can potentially be useful in assessing the accuracy of contextual gene annotation based on such interaction networks.

Computational Biology↗

The CATH Dictionary of Homologous Superfamilies (DHS): a consensus approach for identifying distant structural homologues.

A consensus approach has been developed for identifying distant structural homologues. This is based on the CATH Dictionary of Homologous Superfamilies (DHS), a database of validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies (URL: http://www. biochem.ucl.ac.uk/bsm/dhs). Multiple structural alignments have been generated for 362 well-populated superfamilies in the CATH structural domain database and annotated with secondary structure, physicochemical properties, functional sequence patterns and protein-ligand interaction data. Consensus functional information for each superfamily includes descriptions and keywords extracted from SWISS-PROT and the ENZYME database. The Dictionary provides a powerful resource to validate, examine and visualize key structural and functional features of each homologous superfamily. The value of the DHS, for assessing functional variability and identifying distant evolutionary relationships, is illustrated using the pyridoxal-5'-phosphate (PLP) binding aspartate aminotransferase superfamily. The DHS also provides a tool for examining sequence-structure relationships for proteins within each fold group.

Amino Acid Sequence↗

Correlation between gene expression and GO semantic similarity.

This research analyzes some aspects of the relationship between gene expression, gene function, and gene annotation. Many recent studies are implicitly based on the assumption that gene products that are biologically and functionally related would maintain this similarity both in their expression profiles as well as in their Gene Ontology (GO) annotation. We analyze how accurate this assumption proves to be using real publicly available data. We also aim to validate a measure of semantic similarity for GO annotation. We use the Pearson correlation coefficient and its absolute value as a measure of similarity between expression profiles of gene products. We explore a number of semantic similarity measures (Resnik, Jiang, and Lin) and compute the similarity between gene products annotated using the GO. Finally, we compute correlation coefficients to compare gene expression similarity against GO semantic similarity. Our results suggest that the Resnik similarity measure outperforms the others and seems better suited for use in Gene Ontology. We also deduce that there seems to be correlation between semantic similarity in the GO annotation and gene expression for the three GO ontologies. We show that this correlation is negligible up to a certain semantic similarity value; then, for higher similarity values, the relationship trend becomes almost linear. These results can be used to augment the knowledge provided by clustering algorithms and in the development of bioinformatic tools for finding and characterizing gene products.

Algorithms↗

How well do we understand the clusters found in microarray data?

We wished to quantify the state-of-the-art of our understanding of clusters in microarray data. To do this we systematically compared the clusters produced on sets of microarray data using a representative set of clustering algorithms (hierarchical, k-means, and a modified version of QT_CLUST) with the annotation schemes MIPS, GeneOntology and GenProtEC. We assumed that if a cluster reflected known biology its members would share related ontological annotations. This assumption is the basis of "guilt-by-association" and is commonly used to assign the putative function of proteins. To statistically measure the relationship between cluster and annotation we developed a new predictive discriminatory measure. We found that the clusters found in microarray data do not in general agree with functional annotation classes. Although many statistically significant relationships can be found, the majority of clusters are not related to known biology (as described in annotation ontologies). This implies that use of guilt-by-association is not supported by annotation ontologies. Depending on the estimate of the amount of noise in the data, our results suggest that bioinformatics has only codified a small proportion of the biological knowledge required to understand microarray data.

Algorithms↗

The SWISS-MODEL Repository: new features and functionalities.

The SWISS-MODEL Repository is a database of annotated 3D protein structure models generated by the SWISS-MODEL homology-modelling pipeline. As of September 2005, the repository contained 675,000 models for 604,000 different protein sequences of the UniProt database. Regular updates ensure that the content of the repository reflects the current state of sequence and structure databases, integrating new or modified target sequences, and making use of new template structures. Each Repository entry consists of one or more 3D models accompanied by detailed information about the target protein and the model building process: functional annotation, a detailed template selection log, target-template alignment, summary of the model building and model quality assessment. The SWISS-MODEL Repository is freely accessible at http://swissmodel.expasy.org/repository/.

Computer Graphics↗

Phylogenetic detection of conserved gene clusters in microbial genomes.

BACKGROUND: Microbial genomes contain an abundance of genes with conserved proximity forming clusters on the chromosome. However, the conservation can be a result of many factors such as vertical inheritance, or functional selection. Thus, identification of conserved gene clusters that are under functional selection provides an effective channel for gene annotation, microarray screening, and pathway reconstruction. The problem of devising a robust method to identify these conserved gene clusters and to evaluate the significance of the conservation in multiple genomes has a number of implications for comparative, evolutionary and functional genomics as well as synthetic biology. RESULTS: In this paper we describe a new method for detecting conserved gene clusters that incorporates the information captured by a genome phylogenetic tree. We show that our method can overcome the common problem of overestimation of significance due to the bias in the genome database and thereby achieve better accuracy when detecting functionally connected gene clusters. Our results can be accessed at database GeneChords http://genomics10.bu.edu/GeneChords. CONCLUSION: The methodology described in this paper gives a scalable framework for discovering conserved gene clusters in microbial genomes. It serves as a platform for many other functional genomic analyses in microorganisms, such as operon prediction, regulatory site prediction, functional annotation of genes, evolutionary origin and development of gene clusters.

Algorithms↗

Computationally analyzing the possible biological function of YJL103C--an ORF potentially involved in the regulation of energy process in yeast.

Although the complete genomes of a number of organisms have been sequenced, the biological functions of many genes are still not known. Because experimentally studying the functions of those genes one by one requires tremendous time, it is vital to use published resources like microarray gene expression data for computational analysis of gene functions. One example is YJL103C, a yeast gene of unknown function in the Saccharomyces Genome Database (SGD). It is possible to quickly infer its biological function by computational analysis. In this study, we present an efficient model to explore the biological function of a novel gene using microarray data. We showed that the expression pattern of YJL103C is most similar to the genes in the energy group and respiratory chain subgroup. We further found that YJL103C contains a HAP2,3,4 box in its promoter region and a cytochrome C heme-binding signature in its protein sequence. Our findings define a potential role for YJL103C in the regulation of energy metabolism, specifically in the process of oxidative phosphorylation. Similar bioinformatics methods can be applied to infer the biological functions of other novel genes in organisms for which microarray data are available. In this work, we selected a single gene of unknown function as a case study. By focusing on the power of computer analysis and bioinformatics on the available microarray data, we have determined the likely biological function of YJL103C. Our study provides a method by which to explore the potential function of other genes currently annotated as having an unknown function in any organism for which global gene expression data are available.

Binding Sites↗

PipeOnline 2.0: automated EST processing and functional data sorting.

Expressed sequence tags (ESTs) are generated and deposited in the public domain, as redundant, unannotated, single-pass reactions, with virtually no biological content. PipeOnline automatically analyses and transforms large collections of raw DNA-sequence data from chromatograms or FASTA files by calling the quality of bases, screening and removing vector sequences, assembling and rewriting consensus sequences of redundant input files into a unigene EST data set and finally through translation, amino acid sequence similarity searches, annotation of public databases and functional data. PipeOnline generates an annotated database, retaining the processed unigene sequence, clone/file history, alignments with similar sequences, and proposed functional classification, if available. Functional annotation is automatic and based on a novel method that relies on homology of amino acid sequence multiplicity within GenBank records. Records are examined through a function ordered browser or keyword queries with automated export of results. PipeOnline offers customization for individual projects (MyPipeOnline), automated updating and alert service. PipeOnline is available at http://stress-genomics.org.

Automation↗

Identification of similar regions of protein structures using integrated sequence and structure analysis tools.

BACKGROUND: Understanding protein function from its structure is a challenging problem. Sequence based approaches for finding homology have broad use for annotation of both structure and function. 3D structural information of protein domains and their interactions provide a complementary view to structure function relationships to sequence information. We have developed a web site http://www.sblest.org/ and an API of web services that enables users to submit protein structures and identify statistically significant neighbors and the underlying structural environments that make that match using a suite of sequence and structure analysis tools. To do this, we have integrated S-BLEST, PSI-BLAST and HMMer based superfamily predictions to give a unique integrated view to prediction of SCOP superfamilies, EC number, and GO term, as well as identification of the protein structural environments that are associated with that prediction. Additionally, we have extended UCSF Chimera and PyMOL to support our web services, so that users can characterize their own proteins of interest. RESULTS: Users are able to submit their own queries or use a structure already in the PDB. Currently the databases that a user can query include the popular structural datasets ASTRAL 40 v1.69, ASTRAL 95 v1.69, CLUSTER50, CLUSTER70 and CLUSTER90 and PDBSELECT25. The results can be downloaded directly from the site and include function prediction, analysis of the most conserved environments and automated annotation of query proteins. These results reflect both the hits found with PSI-BLAST, HMMer and with S-BLEST. We have evaluated how well annotation transfer can be performed on SCOP ID's, Gene Ontology (GO) ID's and EC Numbers. The method is very efficient and totally automated, generally taking around fifteen minutes for a 400 residue protein. CONCLUSION: With structural genomics initiatives determining structures with little, if any, functional characterization, development of protein structure and function analysis tools are a necessary endeavor. We have developed a useful application towards a solution to this problem using common structural and sequence based analysis tools. These approaches are able to find statistically significant environments in a database of protein structure, and the method is able to quantify how closely associated each environment is to a predicted functional annotation.

Computational Biology↗

Development and characterization of a normalized canine retinal cDNA library for genomic and expression studies.

PURPOSE: Identification of causative mutations for retinal blinding disorders is often limited by restricted understanding of gene expression and underlying molecular mechanisms that trigger degenerative processes. This study was conducted to develop a catalog of canine retina-expressed genes that would provide a unique tool to investigate normal and altered function in the adult retina. Because of the conserved syntenies between the dog and human, this approach would identify new potential disease candidate genes for both species. METHODS: A canine normalized retinal cDNA library was produced and analyzed by using a modified PhredPhrap algorithm. Computerized annotation provided gene homology and chromosomal location for individual clones and contigs in a Web-accessible database. RESULTS: From 6316 cDNA clones, 3980 retinal expressed sequence tags (ESTs) were derived. Homology to the canine genome draft sequence was found for more than 99% of all ESTs, but only for 32% when compared with annotated canine cDNAs. Functional analysis suggests an enrichment of this library for genes involved with eye function and development, chaperone, or ribosomal functions when compared with mouse and human National Center for Biotechnology Information (NCBI) RefSeq entries. CONCLUSIONS: A combination of annotation approaches with ongoing mapping and expression studies provide functional data covering at least 27% to 30% of the currently proposed canine catalog of genes expressed in the retina. This is an essential first step toward establishing an integrated network for gene identification and expression patterns suitable for functional genetics, comparative genomics and evolutionary analysis of genes and gene families with respect to the developmental and degenerative processes of the retina.

Animals↗

Evolution of enzyme superfamilies.

Enzyme evolution is often constrained by aspects of catalysis. Sets of homologous proteins that catalyze different overall reactions but share an aspect of catalysis, such as a common partial reaction, are called mechanistically diverse superfamilies. The common mechanistic steps and structural characteristics of several of these superfamilies, including the enolase, Nudix, amidohydrolase, and haloacid dehalogenase superfamilies have been characterized. In addition, studies of mechanistically diverse superfamilies are helping to elucidate mechanisms of functional diversification, such as catalytic promiscuity. Understanding how enzyme superfamilies evolve is vital for accurate genome annotation, predicting protein functions, and protein engineering.

Binding Sites↗

MICheck: a web tool for fast checking of syntactic annotations of bacterial genomes.

The annotation of newly sequenced bacterial genomes begins with running several automatic analysis methods, with major emphasis on the identification of protein-coding genes. DNA sequences are heterogeneous in local nucleotide composition and this leads sometimes to sequences being annotated as authentic genes when they are not protein-coding genes or are true but uncharacterized protein-coding genes. This first annotation step is generally followed by an expert manual annotation of the predicted genes. The genomic data (sequence and annotations) organized in an appropriate databank file format is subsequently submitted to an entry point of the International Nucleotide Sequence Database. These procedures are inevitably subject to mistakes, and this can lead to unintentional syntactic annotation errors being stored in public databanks. Here, we present a new web program, MICheck (MIcrobial genome Checker), that enables rapid verification of sets of annotated genes and frameshifts in previously published bacterial genomes. The web interface allows one easily to investigate the MICheck results, i.e. inaccurate or missed gene annotations: a graphical representation is drawn, in which the genomic context of a unique coding DNA sequence annotation or a predicted frameshift is given, using information on the coding potential (curves) and annotation of the neighbouring genes. We illustrate some capabilities of the MICheck site through the analysis of 20 bacterial genomes, 9 of which were selected for their 'Reviewed' status in the National Center for Biotechnology Information (NCBI) Reference Sequence Project (RefSeq). In the context of the numerous re-annotation projects for microbial genomes, this tool can be seen as a preliminary step before the functional re-annotation step to check quickly for missing or wrongly annotated genes. The MICheck website is accessible at the following address: http://www.genoscope.cns.fr/agc/tools/micheck.

Computer Graphics↗

Rhodopseudomonas palustris regulons detected by cross-species analysis of alphaproteobacterial genomes.

Rhodopseudomonas palustris, an alpha-proteobacterium, carries out three of the chemical reactions that support life on this planet: the conversion of sunlight to chemical-potential energy; the absorption of carbon dioxide, which it converts to cellular material; and the fixation of atmospheric nitrogen into ammonia. Insight into the transcription-regulatory network that coordinates these processes is fundamental to understanding the biology of this versatile bacterium. With this goal in mind, we predicted regulatory signals genomewide, using a two-step phylogenetic-footprinting and clustering process that we had developed previously. In the first step, 4,963 putative transcription factor binding sites, upstream of 2,044 genes and operons, were identified using cross-species Gibbs sampling. Bayesian motif clustering was then employed to group the cross-species motifs into regulons. We have identified 101 putative regulons in R. palustris, including 8 that are of particular interest: a photosynthetic regulon, a flagellar regulon, an organic hydroperoxide resistance regulon, the LexA regulon, and four regulons related to nitrogen metabolism (FixK2, NnrR, NtrC, and sigma54). In some cases, clustering allowed us to assign functions to proteins that previously had been annotated with only putative functions; we have identified RPA0828 as the organic hydroperoxide resistance regulator and RPA1026 as a cell cycle methylase. In addition to predicting regulons, we identified a novel inverted repeat that likely forms a highly conserved stem-loop and that occurs downstream of over 100 genes.

Alphaproteobacteria↗

Integrating 'omic' information: a bridge between genomics and systems biology.

The availability of genome sequences for several organisms, including humans, and the resulting first-approximation lists of genes, have allowed a transition from molecular biology to 'modular biology'. In modular biology, biological processes of interest, or modules, are studied as complex systems of functionally interacting macromolecules. Functional genomic and proteomic ('omic') approaches can be helpful to accelerate the identification of the genes and gene products involved in particular modules, and to describe the functional relationships between them. However, the data emerging from individual omic approaches should be viewed with caution because of the occurrence of false-negative and false-positive results and because single annotations are not sufficient for an understanding of gene function. To increase the reliability of gene function annotation, multiple independent datasets need to be integrated. Here, we review the recent development of strategies for such integration and we argue that these will be important for a systems approach to modular biology.

Computational Biology↗