Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Evaluation of features for catalytic residue prediction in novel folds.

Structural genomics projects are determining the three-dimensional structure of proteins without full characterization of their function. A critical part of the annotation process involves appropriate knowledge representation and prediction of functionally important residue environments. We have developed a method to extract features from sequence, sequence alignments, three-dimensional structure, and structural environment conservation, and used support vector machines to annotate homologous and nonhomologous residue positions based on a specific training set of residue functions. In order to evaluate this pipeline for automated protein annotation, we applied it to the challenging problem of prediction of catalytic residues in enzymes. We also ranked the features based on their ability to discriminate catalytic from noncatalytic residues. When applying our method to a well-annotated set of protein structures, we found that top-ranked features were a measure of sequence conservation, a measure of structural conservation, a degree of uniqueness of a residue's structural environment, solvent accessibility, and residue hydrophobicity. We also found that features based on structural conservation were complementary to those based on sequence conservation and that they were capable of increasing predictor performance. Using a family nonredundant version of the ASTRAL 40 v1.65 data set, we estimated that the true catalytic residues were correctly predicted in 57.0% of the cases, with a precision of 18.5%. When testing on proteins containing novel folds not used in training, the best features were highly correlated with the training on families, thus validating the approach to nonhomologous catalytic residue prediction in general. We then applied the method to 2781 coordinate files from the structural genomics target pipeline and identified both highly ranked and highly clustered groups of predicted catalytic residues.

Algorithms↗

Structures of phosphate and trivanadate complexes of Bacillus stearothermophilus phosphatase PhoE: structural and functional analysis in the cofactor-dependent phosphoglycerate mutase superfamily.

Bacillus stearothermophilus phosphatase PhoE is a member of the cofactor-dependent phosphoglycerate mutase superfamily possessing broad specificity phosphatase activity. Its previous structural determination in complex with glycerol revealed probable bases for its efficient hydrolysis of both large, hydrophobic, and smaller, hydrophilic substrates. Here we report two further structures of PhoE complexes, to higher resolution of diffraction, which yield a better and thorough understanding of its catalytic mechanism. The environment of the phosphate ion in the catalytic site of the first complex strongly suggests an acid-base catalytic function for Glu83. It also reveals how the C-terminal tail ordering is linked to enzyme activation on phosphate binding by a different mechanism to that seen in Escherichia coli phosphoglycerate mutase. The second complex structure with an unusual doubly covalently bound trivanadate shows how covalent modification of the phosphorylable His10 is accompanied by small structural changes, presumably to catalytic advantage. When compared with structures of related proteins in the cofactor-dependent phosphoglycerate mutase superfamily, an additional phosphate ligand, Gln22, is observed in PhoE. Functional constraints lead to the corresponding residue being conserved as Gly in fructose-2,6-bisphosphatases and Thr/Ser/Cys in phosphoglycerate mutases. A number of sequence annotation errors in databases are highlighted by this analysis. B. stearothermophilus PhoE is evolutionarily related to a group of enzymes primarily present in Gram-positive bacilli. Even within this group substrate specificity is clearly variable highlighting the difficulties of computational functional annotation in the cofactor-dependent phosphoglycerate mutase superfamily.

Amino Acid Sequence↗

Assigning function to CDS through qualified query answering: beyond alignment and motifs.

In this paper, we show how to use qualitative query answering to annotate CDS-to-function relationships with confidence in the score, confidence in the tool, and confidence in the decision about the function. The system, implemented in Prolog, provides users with a powerful tool to analyze large quantities of data that have been produce by multiple sequence analysis programs. Using qualified query answering techniques, users can easily change the criteria for how tools reinforce each other and for how numbers of occurrences of particular functions reinforce each other. They can also alter how different scores for different tools are categorized.

Animals↗

Effects of resveratrol on gene expression in renal cell carcinoma.

Studies have shown that Resveratrol (RE) can inhibit cancer initiation, promotion, and progression. However the gene expression profile in renal cell carcinoma (RCC) in response to RE treatment has never been reported. To understand the potential anticancer effect of RE on RCC at molecular level, we profiled and analyzed the expression of 2059 cancer-related genes in a RCC cell line RCC54 treated with RE. Biological functions of 633 genes were annotated based on biological process ontology and clustered into functional categories. Twenty-nine highly differentially expressed genes in RE treated RCC54 were identified and the potential implications of some gene expression alterations in RCC carcinogenesis were identified. RE was also shown to inhibit cell growth and induce cell death of RCC cells. The expression alterations of selected genes were validated using reverse transcription polymerase chain reaction. In addition, the gene expression profiles under different RE treatments were analyzed and visualized using singular value decomposition. The findings from this study support the hypothesis that RE induces differential expression of genes that are directly or indirectly related to the inhibition of RCC cell growth and induction of RCC cell death. In addition, it is apparent that the gene expression alterations due to RE treatment depend strongly on RE concentration. This study provides a general understanding of the overall genetic response of RCC54 to RE treatment and yields insights into the understanding of the cancer preventive mechanism of RE in RCC.

Angiogenesis Inhibitors↗

Efficient recognition of protein fold at low sequence identity by conservative application of Psi-BLAST: validation.

A substantial fraction of protein sequences derived from genomic analyses is currently classified as representing 'hypothetical proteins of unknown function'. In part, this reflects the limitations of methods for comparison of sequences with very low identity. We evaluated the effectiveness of a Psi-BLAST search strategy to identify proteins of similar fold at low sequence identity. Psi-BLAST searches for structurally characterized low-sequence-identity matches were carried out on a set of over 300 proteins of known structure. Searches were conducted in NCBI's non-redundant database and were limited to three rounds. Some 614 potential homologs with 25% or lower sequence identity to 166 members of the search set were obtained. Disregarding the expect value, level of sequence identity and span of alignment, correspondence of fold between the target and potential homolog was found in more than 95% of the Psi-BLAST matches. Restrictions on expect value or span of alignment improved the false positive rate at the expense of eliminating many true homologs. Approximately three-quarters of the putative homologs obtained by three rounds of Psi-BLAST revealed no significant sequence similarity to the target protein upon direct sequence comparison by BLAST, and therefore could not be found by a conventional search. Although three rounds of Psi-BLAST identified many more homologs than a standard BLAST search, most homologs were undetected. It appears that more than 80% of all homologs to a target protein may be characterized by a lack of significant sequence similarity. We suggest that conservative use of Psi-BLAST has the potential to propose experimentally testable functions for the majority of proteins currently annotated as 'hypothetical proteins of unknown function'.

Algorithms↗

In silico predictions of Escherichia coli metabolic capabilities are consistent with experimental data.

A significant goal in the post-genome era is to relate the annotated genome sequence to the physiological functions of a cell. Working from the annotated genome sequence, as well as biochemical and physiological information, it is possible to reconstruct complete metabolic networks. Furthermore, computational methods have been developed to interpret and predict the optimal performance of a metabolic network under a range of growth conditions. We have tested the hypothesis that Escherichia coli uses its metabolism to grow at a maximal rate using the E. coli MG1655 metabolic reconstruction. Based on this hypothesis, we formulated experiments that describe the quantitative relationship between a primary carbon source (acetate or succinate) uptake rate, oxygen uptake rate, and maximal cellular growth rate. We found that the experimental data were consistent with the stated hypothesis, namely that the E. coli metabolic network is optimized to maximize growth under the experimental conditions considered. This study thus demonstrates how the combination of in silico and experimental biology can be used to obtain a quantitative genotype-phenotype relationship for metabolism in bacterial cells.

Acetates↗

RetrOryza: a database of the rice LTR-retrotransposons.

Long terminal repeat (LTR)-retrotransposons comprise a significant portion of the rice genome. Their complete characterization is thus necessary if the sequenced genome is to be annotated correctly. In addition, because LTR-retrotransposons can influence the expression of neighboring genes, the complete identification of these elements in the rice genome is essential in order to study their putative functional interactions with the plant genes. The aims of the database are to (i) Assemble a comprehensive dataset of LTR-retrotransposons that includes not only abundant elements, but also low copy number elements. (ii) Provide an interface to efficiently access the resources stored in the database. This interface should also allow the community to annotate these elements. (iii) Provide a means for identifying LTR-retrotransposons inserted near genes. Here we present the results, where 242 complete LTR-retrotransposons have been structurally and functionally annotated. A web interface to the database has been made available (http://www.retroryza.org/), through which the user can annotate a sequence or search for LTR-retrotransposons in the neighborhood of a gene of interest.

Databases, Nucleic Acid↗

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors↗

Emergence of a subfamily of xylanase inhibitors within glycoside hydrolase family 18.

The xylanase inhibitor protein I (XIP-I), recently identified in wheat, inhibits xylanases belonging to glycoside hydrolase families 10 (GH10) and 11 (GH11). Sequence and structural similarities indicate that XIP-I is related to chitinases of family GH18, despite its lack of enzymatic activity. Here we report the identification and biochemical characterization of a XIP-type inhibitor from rice. Despite its initial classification as a chitinase, the rice inhibitor does not exhibit chitinolytic activity but shows specificities towards fungal GH11 xylanases similar to that of its wheat counterpart. This, together, with an analysis of approximately 150 plant members of glycosidase family GH18 provides compelling evidence that xylanase inhibitors are largely represented in this family, and that this novel function has recently emerged based on a common scaffold. The plurifunctionality of GH18 members has major implications for genomic annotations and predicted gene function. This study provides new information which will lead to a better understanding of the biological significance of a number of GH18 'inactivated' chitinases.

Amino Acid Sequence↗

Protein expression, crystallization and preliminary X-ray crystallographic studies of YjbK from Bacillus subtilis.

B. subtilis YjbK is a protein with 190 residues of uncharacterized function, it has been annotated by Pfam database as a member of adenylate cyclase family (EC: 4.6.1.1). In order to identify its exact function via structural studies, yjbK gene was amplified from B. subtilis genomic DNA and cloned into expression vector pET21-DEST. The protein was expressed in a soluble form in E. coli and purified to homogeneity. YjbK was crystallized and diffracted to a resolution of 2.0 A in-house. The crystals belong to P1 space group, with unit cell parameters a = 32.38 A, b = 34.69 A, c = 46.02 A, alpha = 96.560 degrees, beta = 99.683 degrees, gamma = 111.333 degrees. There is one molecule per asymmetric unit.

Amino Acid Sequence↗

Finding GeneRIFs via gene ontology annotations.

A Gene Reference Into Function (GeneRIF) is a concise phrase describing a function of a gene in the Entrez Gene database. Applying techniques from the area of natural language processing known as automatic summarization, it is possible to link the Entrez Gene database, the Gene Ontology, and the biomedical literature. A system was implemented that automatically suggests a sentence from a PubMed/MEDLINE abstract as a candidate GeneRIF by exploiting a gene's GO annotations along with location features and cue words. Results suggest that the method can significantly increase the number of GeneRIF annotations in Entrez Gene, and that it produces qualitatively more useful GeneRIFs than other methods.

Algorithms↗

[Gene regulation and bioinformatics].

Gene regulation networks control differentiation and function of hundreds of cell types. Dysfunctions of transcription factors, which are key elements in the regulation pathways, are involved in numerous pathologies. The recent development of genomics technologies allows the study of gene regulation mechanisms and help us understand their impact on the cells. Bioinformatic tools are needed to fully exploit data obtained by genomics approaches. Thus, bioinformatics play an essential role in the characterisation of the transcription factors binding sites and their target genes. In this review we will introduce the main breakthroughs in bioinformatics area for the comprehension of regulation mechanisms. We will insist on i) the approaches "with a priori" for the genome annotation based on known transcription factors binding sites, ii) the approaches "without a priori" for the discovery of new binding sites and iii) the functional annotation of the target genes of those transcription factors. Finally, we will present recent examples of the fruitful use of in silico studies for the comprehension of regulation mechanisms and of the consequences of their dysfunction.

Algorithms↗

MitoP2: the mitochondrial proteome database--now including mouse data.

The MitoP2 database (http://www.mitop.de) integrates information on mitochondrial proteins, their molecular functions and associated diseases. The central database features are manually annotated reference proteins localized or functionally associated with mitochondria supplied for yeast, human and mouse. MitoP2 enables (i) the identification of putative orthologous proteins between these species to study evolutionarily conserved functions and pathways; (ii) the integration of data from systematic genome-wide studies such as proteomics and deletion phenotype screening; (iii) the prediction of novel mitochondrial proteins using data integration and the assignment of evidence scores; and (iv) systematic searches that aim to find the genes that underlie common and rare mitochondrial diseases. The data and analysis files are referenced to data sources in PubMed and other online databases and can be easily downloaded. MitoP2 users can explore the relationship between mitochondrial dysfunctions and disease and utilize this information to conduct systems biology approaches on mitochondria.

Animals↗

Plant Gene and Alternatively Spliced Variant Annotator. A plant genome annotation pipeline for rice gene and alternatively spliced variant identification with cross-species expressed sequence tag conservation from seven plant species.

The completion of the rice (Oryza sativa) genome draft has brought unprecedented opportunities for genomic studies of the world's most important food crop. Previous rice gene annotations have relied mainly on ab initio methods, which usually yield a high rate of false-positive predictions and give only limited information regarding alternative splicing in rice genes. Comparative approaches based on expressed sequence tags (ESTs) can compensate for the drawbacks of ab initio methods because they can simultaneously identify experimental data-supported genes and alternatively spliced transcripts. Furthermore, cross-species EST information can be used to not only offset the insufficiency of same-species ESTs but also derive evolutionary implications. In this study, we used ESTs from seven plant species, rice, wheat (Triticum aestivum), maize (Zea mays), barley (Hordeum vulgare), sorghum (Sorghum bicolor), soybean (Glycine max), and Arabidopsis (Arabidopsis thaliana), to annotate the rice genome. We developed a plant genome annotation pipeline, Plant Gene and Alternatively Spliced Variant Annotator (PGAA). Using this approach, we identified 852 genes (931 isoforms) not annotated in other widely used databases (i.e. the Institute for Genomic Research, National Center for Biotechnology Information, and Rice Annotation Project) and found 87% of them supported by both rice and nonrice EST evidence. PGAA also identified more than 44,000 alternatively spliced events, of which approximately 20% are not observed in the other three annotations. These novel annotations represent rich opportunities for rice genome research, because the functions of most of our annotated genes are currently unknown. Also, in the PGAA annotation, the isoforms with non-rice-EST-supported exons are significantly enriched in transporter activity but significantly underrepresented in transcription regulator activity. We have also identified potential lineage-specific and conserved isoforms, which are important markers in evolutionary studies. The data and the Web-based interface, RiceViewer, are available for public access at http://RiceViewer.genomics.sinica.edu.tw/.

Base Sequence↗

The TIGR gene indices: reconstruction and representation of expressed gene sequences.

Expressed sequence tags (ESTs) have provided a first glimpse of the collection of transcribed sequences in a variety of organisms. However, a careful analysis of this sequence data can provide significant additional functional, structural and evolutionary information. Our analysis of the public EST sequences, available through the TIGR Gene Indices (TGI; http://www.tigr.org/tdb/tdb.html ), is an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed for selected organisms by first clustering, then assembling EST and annotated gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, and to provide links between orthologous and paralogous genes.

Base Sequence↗

Predicting gene function from gene expressions and ontologies.

We introduce a methodology for inducing predictive rule models for functional classification of gene expressions from microarray hybridisation experiments. The basic learning method is the rough set framework for rule induction. The methodology is different from the commonly used unsupervised clustering approaches in that it exploits background knowledge of gene function in a supervised manner. Genes are annotated using Ashburner's Gene Ontology and the functional classes used for learning are mined from these annotations. From the original expression data, we extract a set of biologically meaningful features that are used for learning. A rule model is induced from the data described in terms of these features. Its predictive quality is fine-turned via cross-validation on subsets of the known genes prior to classification of unknown genes. The predictive and descriptive quality of such a rule model is demonstrated on the fibroblast serum response data previously analysed by Iyer et. al. Our analysis shows that the rules are capable of representing the complex relationship between gene expressions and function, and that it is possible to put forward high quality hypotheses about the function of unknown genes.

Algorithms↗

DNA transposons in vertebrate functional genomics.

Genome sequences of many model organisms of developmental or agricultural importance are becoming available. The tremendous amount of sequence data is fuelling the next phases of challenging research: annotating all genes with functional information, and devising new ways for the experimental manipulation of vertebrate genomes. Transposable elements are known to be efficient carriers of foreign DNA into cells. Notably, members of the Tc1/mariner and the hAT transposon families retain their high transpositional activities in species other than their hosts. Indeed, several of these elements have been successfully used for transgenesis and insertional mutagenesis, expanding our abilities in genome manipulations in vertebrate model organisms. Transposon-based genetic tools can help scientists to understand mechanisms of embryonic development and pathogenesis, and will likely contribute to successful human gene therapy. We discuss the possibilities of transposon-based techniques in functional genomics, and review the latest results achieved by the most active DNA transposons in vertebrates. We put emphasis on the evolution and regulation of members of the best-characterized and most widely used Tc1/mariner family.

Animals↗