Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Systematic comparison of catalytic mechanisms of hydrolysis and transfer reactions classified in the EzCatDB database.

Catalytic mechanisms of 270 enzymes from 131 superfamilies, mainly hydrolases and transferases, were analyzed based on their enzyme structures. A method of systematic comparison and classification of the catalytic reactions was developed. Hydrolysis and transfer reactions closely resemble one another, displaying common mechanisms, single displacement, and double displacement. These displacement mechanisms might be further subclassified according to the type of catalytic factors and nucleophilic substitution involved. Several types of catalytic factors exist: nucleophile, acid, base, stabilizer, modulator, cofactors. Nucleophilic substitution might be categorized as S(N)1/S(N)2 (or dissociative/associative) reactions. The classification indicates that some mechanisms favor particular types of catalytic factors. In hydrolyses of amide bonds and phosphoric ester bonds, mechanisms with single displacement tend to use inorganic cofactors such as zinc and magnesium ions as important catalysts, whereas those with double displacement frequently do not use such cofactors. In contrast, hydrolyses of O-glycoside bond rarely use such cofactors, with one exception. The trypsin-like hydrolytic reaction, which is catalyzed by the classic catalytic triad comprising serine/histidine/aspartate, can be considered as a "super-reaction" because it is observed in at least three nonhomologous enzymes, whereas most reactions are singlets without any nonhomologous enzymes. By dividing complex reactions into several reactions, correlations between active site structures and catalytic functions can be suggested. This classification method is applicable to other reactions such as elimination and isomerization. Furthermore, it will facilitate annotation of enzyme functions from 3D patterns of enzyme active sites. The classification is available at http://mbs.cbrc.jp/EzCatDB/RLCP/index.html.

Binding Sites↗

Annotation of cis-regulatory elements by identification, subclassification, and functional assessment of multispecies conserved sequences.

An important step toward improving the annotation of the human genome is to identify cis-acting regulatory elements from primary DNA sequence. One approach is to compare sequences from multiple, divergent species. This approach distinguishes multispecies conserved sequences (MCS) in noncoding regions from more rapidly evolving neutral DNA. Here, we have analyzed a region of approximately 238kb containing the human alpha globin cluster that was sequenced and/or annotated across the syntenic region in 22 species spanning 500 million years of evolution. Using a variety of bioinformatic approaches and correlating the results with many aspects of chromosome structure and function in this region, we were able to identify and evaluate the importance of 24 individual MCSs. This approach sensitively and accurately identified previously characterized regulatory elements but also discovered unidentified promoters, exons, splicing, and transcriptional regulatory elements. Together, these studies demonstrate an integrated approach by which to identify, subclassify, and predict the potential importance of MCSs.

Animals↗

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence↗

Mutations in Mycobacterium tuberculosis Rv0444c, the gene encoding anti-SigK, explain high level expression of MPB70 and MPB83 in Mycobacterium bovis.

It has recently been advanced that Mycobacterium tuberculosis sigma factor K (SigK) positively regulates expression of the antigenic proteins MPB70 and MPB83. As expression of these proteins differs between M. tuberculosis (low) and Mycobacterium bovis (high), this study set out to determine whether M. bovis lacks a functional SigK repressor (anti-SigK). By comparing genes near sigK in M. tuberculosis H37Rv and M. bovis AF2122/97, we observed that Rv0444c, annotated as unknown function, had variable sequence in M. bovis. Analysis of in vitro mpt70/mpt83 expression and Rv0444c sequencing across M. tuberculosis complex (MTC) members revealed that high-level expression was associated with a mutated Rv0444c. Complementation of M. bovis bacillus Calmette-Guerin Russia, a high producer of MPB70/MPB83, with wild-type Rv0444c resulted in a significant decrease in mpb70/mpb83 expression. Conversely, a M. tuberculosis H37Rv mutant which expressed sigK but not Rv0444c manifested the M. bovis phenotype of high-level MPB70/MPB83 expression. Further support that Rv0444c encodes the anti-SigK was obtained by yeast two-hybrid studies, where the N-terminal region of Rv0444c-encoded protein interacted with SigK. Together these findings indicate that Rv0444c encodes the regulator of SigK (RskA) and mutations in this gene explain high-level MPT70/MPT83 expression by certain MTC members.

Antigens, Bacterial↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗

Cis-regulatory elements: systematic identification and horticultural applications.

Cis-regulatory elements (CREs) are the genetic DNA fragments bound by transcription factors (TFs). CREs function as molecular switches that precisely modulate the dosage and spatiotemporal patterns of gene expression. The systematic identification of CREs not only facilitates the annotation of the functional non-coding genome but also provides essential insights into the architecture of gene regulatory networks and sheds light on an accurate selection of the target sites for genetic engineering of crops. In this review, we summarize the current high-throughput methodologies used for identifying CREs, illustrate the associations between CREs and agronomic traits in horticultural crops, and discuss how CREs can be exploited to facilitate crop breeding.

Breeding↗

An expression and bioinformatics analysis of the Arabidopsis serine carboxypeptidase-like gene family.

The Arabidopsis (Arabidopsis thaliana) genome encodes a family of 51 proteins that are homologous to known serine carboxypeptidases. Based on their sequences, these serine carboxypeptidase-like (SCPL) proteins can be divided into several major clades. The first group consists of 21 proteins which, despite the function implied by their annotation, includes two that have been shown to function as acyltransferases in plant secondary metabolism: sinapoylglucose:malate sinapoyltransferase and sinapoylglucose:choline sinapoyltransferase. A second group comprises 25 SCPL proteins whose biochemical functions have not been clearly defined. Genes encoding representatives from both of these clades can be found in many plants, but have not yet been identified in other phyla. In contrast, the remaining SCPL proteins include five members that are similar to serine carboxypeptidases from a variety of organisms, including fungi and animals. Reverse transcription PCR results suggest that some SCPL genes are expressed in a highly tissue-specific fashion, whereas others are transcribed in a wide range of tissue types. Taken together, these data suggest that the Arabidopsis SCPL gene family encodes a diverse group of enzymes whose functions are likely to extend beyond protein degradation and processing to include activities such as the production of secondary metabolites.

Arabidopsis↗

Smc01944, a secreted peroxidase induced by oxidative stresses in Sinorhizobium meliloti 1021.

Sequencing of the Sinorhizobium meliloti strain 1021 genome led to the detection of 6204 open reading frames, 41 % of which have no hypothetical function. To help annotate this genome, a transcriptome analysis was carried out with a dedicated microarray consisting of 146 genes belonging to three different classes: (i) no hypothetical function; (ii) potentially involved in oxidative stress responses; (iii) known to participate in oxidative stress responses (e.g. catalase and superoxide dismutase genes). This transcriptome analysis, together with biological experiments and in silico investigations, identified new genes induced by exogenous H(2)O(2). The smc01944 gene was the most strongly induced: quantitative PCR showed that the amount of smc01944 mRNA increased 50-fold following the addition of 10 mM H(2)O(2), whereas the amount of katA mRNA (encoding a catalase) only increased 10-fold. Smc01944 is a non-haem chloroperoxidase (Cpo). The only member of this family to have been so far characterized is encoded by prxC of Pseudomonas fluorescens. Unexpectedly, the NH(2)-terminus of Smc01944 includes a signal peptide and Smc01944 is secreted into the supernatant. Interestingly, smc01944 is preceded by smc01945, encoding an OhrR-like regulator (MarR family). Thus, Smc01944 is the first exported Cpo encoded by a gene possibly regulated by an OhrR regulator. It was also shown that smc01944 is induced by t-butyl and cumene hydroperoxides but only slightly by menadione. The study of Smc01944 described in this work showed that the oxidative stress response of S. meliloti seems to differ from that of other bacteria characterized to date.

Amino Acid Sequence↗

PRIME: automatically extracted PRotein Interactions and Molecular Information databasE.

With the exponentially increasing amount of information in the biomedical field, the significance of advanced information retrieval and information extraction, as well as the role of databases, has been increasing. PRIME is an integrated gene/protein informatics database based on natural language processing. It provides automatically extracted protein/family/gene/compound interaction information including both physical and genetic interactions, gene ontology based functions, and graphic pathway viewers. Gene/protein/family names and functional terms are recognized based on dictionaries developed in our laboratory. The interaction and functional information are extracted by syntactic dependencies and various phrase patterns. We have included about 920,000 (non-redundant) protein interactions and 360,000 annotated gene-function relationships for major eukaryotes. By combining the sequence and text information, the pathway comparison between two organisms and simple pathway deduction based on other organism interaction data, and pathway filtering using tissue expression data, are also available. This database is accessible at http://prime.ontology.ims.u-tokyo.ac.jp:8081.

Abstracting and Indexing↗

Towards precise classification of cancers based on robust gene functional expression profiles.

BACKGROUND: Development of robust and efficient methods for analyzing and interpreting high dimension gene expression profiles continues to be a focus in computational biology. The accumulated experiment evidence supports the assumption that genes express and perform their functions in modular fashions in cells. Therefore, there is an open space for development of the timely and relevant computational algorithms that use robust functional expression profiles towards precise classification of complex human diseases at the modular level. RESULTS: Inspired by the insight that genes act as a module to carry out a highly integrated cellular function, we thus define a low dimension functional expression profile for data reduction. After annotating each individual gene to functional categories defined in a proper gene function classification system such as Gene Ontology applied in this study, we identify those functional categories enriched with differentially expressed genes. For each functional category or functional module, we compute a summary measure (s) for the raw expression values of the annotated genes to capture the overall activity level of the module. In this way, we can treat the gene expressions within a functional module as an integrative data point to replace the multiple values of individual genes. We compare the classification performance of decision trees based on functional expression profiles with the conventional gene expression profiles using four publicly available datasets, which indicates that precise classification of tumour types and improved interpretation can be achieved with the reduced functional expression profiles. CONCLUSION: This modular approach is demonstrated to be a powerful alternative approach to analyzing high dimension microarray data and is robust to high measurement noise and intrinsic biological variance inherent in microarray data. Furthermore, efficient integration with current biological knowledge has facilitated the interpretation of the underlying molecular mechanisms for complex human diseases at the modular level.

Algorithms↗

VisANT: data-integrating visual framework for biological networks and modules.

VisANT is a web-based software framework for visualizing and analyzing many types of networks of biological interactions and associations. Networks are a useful computational tool for representing many types of biological data, such as biomolecular interactions, cellular pathways and functional modules. Given user-defined sets of interactions or groupings between genes or proteins, VisANT provides: (i) a visual interface for combining and annotating network data, (ii) supporting function and annotation data for different genomes from the Gene Ontology and KEGG databases and (iii) the statistical and analytical tools needed for extracting topological properties of the user-defined networks. Users can customize, modify, save and share network views with other users, and import basic network data representations from their own data sources, and from standard exchange formats such as PSI-MI and BioPAX. The software framework we employ also supports the development of more sophisticated visualization and analysis functions through its open API for Java-based plug-ins. VisANT is distributed freely via the web at http://visant.bu.edu and can also be downloaded for individual use.

Computer Graphics↗

Annotation in three dimensions. PINTS: Patterns in Non-homologous Tertiary Structures.

The detection of local structural patterns in proteins (e.g. active sites) can provide insights into protein function in the absence of sequence or fold similarity. Methods to detect such similarities are key during structural annotation, for example with results from Structural Genomics initiatives. PINTS (Patterns in Non-homologous Tertiary Structures, http://pints.embl.de) performs database searches for such patterns and most importantly provides a measure of statistical significance for any similarity uncovered. To aid functional annotation of proteins, we allow comparisons of pre-defined patterns against databases of complete structures and of entire structures to databases of particular residues likely to be functionally important.

Binding Sites↗

The Arabidopsis thaliana chloroplast proteome reveals pathway abundance and novel protein functions.

BACKGROUND: Chloroplasts are plant cell organelles of cyanobacterial origin. They perform essential metabolic and biosynthetic functions of global significance, including photosynthesis and amino acid biosynthesis. Most of the proteins that constitute the functional chloroplast are encoded in the nuclear genome and imported into the chloroplast after translation in the cytosol. Since protein targeting is difficult to predict, many nuclear-encoded plastid proteins are still to be discovered. RESULTS: By tandem mass spectrometry, we identified 690 different proteins from purified Arabidopsis chloroplasts. Most proteins could be assigned to known protein complexes and metabolic pathways, but more than 30% of the proteins have unknown functions, and many are not predicted to localize to the chloroplast. Novel structure and function prediction methods provided more informative annotations for proteins of unknown functions. While near-complete protein coverage was accomplished for key chloroplast pathways such as carbon fixation and photosynthesis, fewer proteins were identified from pathways that are downregulated in the light. Parallel RNA profiling revealed a pathway-dependent correlation between transcript and relative protein abundance, suggesting gene regulation at different levels. CONCLUSIONS: The chloroplast proteome contains many proteins that are of unknown function and not predicted to localize to the chloroplast. Expression of nuclear-encoded chloroplast genes is regulated at multiple levels in a pathway-dependent context. The combined shotgun proteomics and RNA profiling approach is of high potential value to predict metabolic pathway prevalence and to define regulatory levels of gene expression on a pathway scale.

Arabidopsis↗

KEGG: kyoto encyclopedia of genes and genomes.

KEGG (Kyoto Encyclopedia of Genes and Genomes) is a knowledge base for systematic analysis of gene functions, linking genomic information with higher order functional information. The genomic information is stored in the GENES database, which is a collection of gene catalogs for all the completely sequenced genomes and some partial genomes with up-to-date annotation of gene functions. The higher order functional information is stored in the PATHWAY database, which contains graphical representations of cellular processes, such as metabolism, membrane transport, signal transduction and cell cycle. The PATHWAY database is supplemented by a set of ortholog group tables for the information about conserved subpathways (pathway motifs), which are often encoded by positionally coupled genes on the chromosome and which are especially useful in predicting gene functions. A third database in KEGG is LIGAND for the information about chemical compounds, enzyme molecules and enzymatic reactions. KEGG provides Java graphics tools for browsing genome maps, comparing two genome maps and manipulating expression maps, as well as computational tools for sequence comparison, graph comparison and path computation. The KEGG databases are daily updated and made freely available (http://www. genome.ad.jp/kegg/).

Animals↗

A functional update of the Escherichia coli K-12 genome.

BACKGROUND: Since the genome of Escherichia coli K-12 was initially annotated in 1997, additional functional information based on biological characterization and functions of sequence-similar proteins has become available. On the basis of this new information, an updated version of the annotated chromosome has been generated. RESULTS: The E. coli K-12 chromosome is currently represented by 4,401 genes encoding 116 RNAs and 4,285 proteins. The boundaries of the genes identified in the GenBank Accession U00096 were used. Some protein-coding sequences are compound and encode multimodular proteins. The coding sequences (CDSs) are represented by modules (protein elements of at least 100 amino acids with biological activity and independent evolutionary history). There are 4,616 identified modules in the 4,285 proteins. Of these, 48.9% have been characterized, 29.5% have an imputed function, 2.1% have a phenotype and 19.5% have no function assignment. Only 7% of the modules appear unique to E. coli, and this number is expected to be reduced as more genome data becomes available. The imputed functions were assigned on the basis of manual evaluation of functions predicted by BLAST and DARWIN analyses and by the MAGPIE genome annotation system. CONCLUSIONS: Much knowledge has been gained about functions encoded by the E. coli K-12 genome since the 1997 annotation was published. The data presented here should be useful for analysis of E. coli gene products as well as gene products encoded by other genomes.

Bacterial Proteins↗

Rank information: a structure-independent measure of evolutionary trace quality that improves identification of protein functional sites.

Protein functional sites are key targets for drug design and protein engineering, but their large-scale experimental characterization remains difficult. The evolutionary trace (ET) is a computational approach to this problem that has been useful in a variety of case studies, but its proteomic scale application is partially hindered because automated retrieval of input sequences from databases often includes some with errors that degrade functional site identification. To recognize and purge these sequences, this study introduces a novel and structure-free measure of ET quality called rank information (RI). It is shown that RI decreases in response to errors in sequences, alignments, or functional classifications. Conversely, an automated procedure to increase RI by selectively removing sequences improves functional site identification so as to nearly match manually curated traces in kinases and in a test set of 79 diverse proteins. Thus we conclude that RI partially reflects the evolutionary consistency of sequence, structure, and function. In practice, as the size of the proteome continues to grow exponentially, it provides a novel and structure-free measure of ET quality that increases its accuracy for large-scale automated annotation of protein functional sites.

Algorithms↗

Enzyme genomics: Application of general enzymatic screens to discover new enzymes.

In all sequenced genomes, a large fraction of predicted genes encodes proteins of unknown biochemical function and up to 15% of the genes with "known" function are mis-annotated. Several global approaches are routinely employed to predict function, including sophisticated sequence analysis, gene expression, protein interaction, and protein structure. In the first coupling of genomics and enzymology, Phizicky and colleagues undertook a screen for specific enzymes using large pools of partially purified proteins and specific enzymatic assays. Here we present an overview of the further developments of this approach, which involve the use of general enzymatic assays to screen individually purified proteins for enzymatic activity. The assays have relaxed substrate specificity and are designed to identify the subclass or sub-subclasses of enzymes (phosphatase, phosphodiesterase/nuclease, protease, esterase, dehydrogenase, and oxidase) to which the unknown protein belongs. Further biochemical characterization of proteins can be facilitated by the application of secondary screens with natural substrates (substrate profiling). We demonstrate here the feasibility and merits of this approach for hydrolases and oxidoreductases, two very broad and important classes of enzymes. Application of general enzymatic screens and substrate profiling can greatly speed up the identification of biochemical function of unknown proteins and the experimental verification of functional predictions produced by other functional genomics approaches.

Enzymes↗

The relationship between protein sequences and their gene ontology functions.

BACKGROUND: One main research challenge in the post-genomic era is to understand the relationship between protein sequences and their biological functions. In recent years, several automated annotation systems have been developed for the functional assignment of uncharacterized proteins. The underlying assumption of these systems is that similar sequences imply similar biological functions. However, it has been noted that matching sequences do not always infer similar functions. RESULTS: In this paper, we present the correlation between protein sequences and protein functions for the yeast proteome in the context of gene ontology. A novel measure is introduced to define the overall similarity between two protein sequences. The effects of the level as well as the size of a gene ontology group on the degree of similarity were studied. The similarity distributions at different levels of gene ontology trees are presented. To evaluate the theoretical prediction power of similar sequences, we computed the posterior probability of correct predictions. CONCLUSION: The results indicate that protein pairs of similar biological functions tend to have higher sequence similarity, although the similarity distribution in each functional group is heterogeneous and varies from group to group. We conclude that sequence similarity can serve as a key measure in protein function prediction. However, the resulting annotations must be verified through other means. A method that combines a broader range of measures is more likely to provide more accurate prediction. Our study indicates that the posterior probability of a correct prediction could serve as one of the key measures.

Amino Acid Sequence↗