Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Enzyme function less conserved than anticipated.

The level of sequence similarity that implies similarity in protein structure is well established. Recently, many groups proposed thresholds for similarity in sequence implying similarity in enzymatic function. All previous results suggest the strong conservation of enzymatic function above levels of 50% pairwise sequence identity. Here, I argue that all groups substantially overestimated the conservation of enzyme function because their data sets were either too biased, or too small. An unbiased analysis suggested that less than 30% of the pair fragments above 50% sequence identity have entirely identical EC numbers. Another surprising finding was that even BLAST E-values below 10(-50) did not suffice to automatically transfer enzyme function without errors. As expected, most misclassifications originated from similarities in relatively short regions and/or from transferring annotations for different domains. Both problems cannot be corrected easily by adjusting the thresholds for automatic transfer of genome annotations. A score relating sequence identity to alignment length (distance from HSSP-threshold) outperformed statistical BLAST scores for high sequence similarity. In particular, the distance score allowed error-free transfer of enzyme function for the 10% most similar enzyme pairs. The results illustrated how difficult it is to assess the conservation of protein function and to guarantee error-free genome annotations, in general: sets with millions of pair comparisons might not suffice to arrive at statistically significant conclusions. In practice, the revised detailed estimates for the sequence conservation of enzyme function may provide important benchmarks for everyday sequence analysis and for more cautious automatic genome annotations.

Amino Acid Sequence↗

Fruits and flies: a genomics perspective of an invertebrate model organism.

The increasing number of species for which a full genome sequence is available offers rich pickings for geneticists, but comparative analysis and assembly of information gathered across species does not always lead to answers about the function of a particular gene. This paper aims to place the invertebrate model system--the fly Drosophila melanogaster--into this playing field and to discuss how the organism arrived at its position in functional genetic analysis. Indeed, despite the wealth of knowledge on how a fly lives, breathes and flies, this organism is likely to remain a player in the analysis of biological, disease and pharmaceutical processes. The fast genetics Drosophila offers, combined with a well-annotated genome and a wealth of techniques facilitating gene function discovery, will ensure its place in functional genomics for some time to come. Although the fly cannot speak, it certainly can tell a tale.

Animals↗

Reverse genetics in the Arabidopsis chloroplast genome identifies rps16 as a transcribed pseudogene.

The plastid (chloroplast) genomes of seed plants contain a conserved set of ribosomal protein genes. The rps16 gene represents an exception: It has been lost from the plastid genomes of gymnosperms and several lineages of angiosperms, and may have undergone pseudogenization in a few other lineages, including members of the Brassicaceae family. Here we report a reverse genetic approach to test the annotated rps16 gene in the Arabidopsis plastid genome for functionality. Employing the recently developed plastid transformation technology for the model plant Arabidopsis, we have deleted the putative rps16 gene from the Arabidopsis plastid genome. We report that the resulting transplastomic plants display wild-type-like growth and photosynthetic performance under a wide range of conditions. Moreover, genome-wide analyses of chloroplast transcript levels and ribosome footprints revealed unaltered plastid translational activity in Δrps16 mutants compared with wild-type plants. We conclude that the annotated rps16 gene in the plastid genome of Arabidopsis is a transcribed pseudogene that has been replaced in evolution by a nuclear gene copy that supplies functional S16 protein to chloroplasts.

Arabidopsis↗

Prediction of yeast protein-protein interaction network: insights from the Gene Ontology and annotations.

A map of protein-protein interactions provides valuable insight into the cellular function and machinery of a proteome. By measuring the similarity between two Gene Ontology (GO) terms with a relative specificity semantic relation, here, we proposed a new method of reconstructing a yeast protein-protein interaction map that is solely based on the GO annotations. The method was validated using high-quality interaction datasets for its effectiveness. Based on a Z-score analysis, a positive dataset and a negative dataset for protein-protein interactions were derived. Moreover, a gold standard positive (GSP) dataset with the highest level of confidence that covered 78% of the high-quality interaction dataset and a gold standard negative (GSN) dataset with the lowest level of confidence were derived. In addition, we assessed four high-throughput experimental interaction datasets using the positives and the negatives as well as GSPs and GSNs. Our predicted network reconstructed from GSPs consists of 40,753 interactions among 2259 proteins, and forms 16 connected components. We mapped all of the MIPS complexes except for homodimers onto the predicted network. As a result, approximately 35% of complexes were identified interconnected. For seven complexes, we also identified some nonmember proteins that may be functionally related to the complexes concerned. This analysis is expected to provide a new approach for predicting the protein-protein interaction maps from other completely sequenced genomes with high-quality GO-based annotations.

Databases, Genetic↗

I-conotoxin superfamily revisited.

The I-conotoxin superfamily (I-Ctx) is known to have four disulfide bonds with the cysteine arrangement C-C-CC-CC-C-C, and the members inhibit or modify ion channels of nerve cells. Recently, Olivera and co-workers (FEBS J. 2005; 272: 4178-4188) have suggested that the previously described I-Ctx should now be divided into two different gene superfamilies, namely, I1 and I2, in view of their having two different types of signal peptides and exhibiting distinct functions. We have revisited the 28 entries presently grouped as I-Ctx in UniProt Swiss-Prot knowledgebase, and on the basis of in silico analysis have divided them into I1 and I2 superfamilies. The sequence analysis has provided a framework for in silico annotation enabling us to carry out computer-based functional characterization of the UniProtKB/TrEMBL entry Q59AA4 from Conus miles and to predict it as a member of the I2 superfamily. Furthermore, we have predicted the mature toxin of this entry and have proposed that it may be an inhibitor of voltage-gated potassium channels.

Amino Acid Sequence↗

A corpus driven approach applying the "frame semantic" method for modeling functional status terminology.

In an effort to unearth semantic models that could prove fruitful to functional-status terminology development we applied the "frame semantic" method, derived from the linguistic theory of thematic roles currently exemplified in the Berkeley "FrameNet" Project. Full descriptive sentences with functional-status conceptual meaning were derived from structured content within a corpus of questionnaire assessment instruments commonly used in clinical practice for functional-status assessment. Syntactic components in those sentences were delineated through manual annotation and mark-up. The annotated syntactic constituents were tagged as frame elements according to their semantic role within the context of the derived functional-status expression. Through this process generalizable "semantic frames" were elaborated with recurring "frame elements". The "frame semantic" method as an approach to rendering semantic models for functional-status terminology development and its use as a basis for machine recognition of functional status data in clinical narratives are discussed.

Information Science↗

The histone database: a comprehensive WWW resource for histones and histone fold-containing proteins.

The Histone Database (HDB) is an annotated and searchable collection of all full-length sequences and structures of histone and non-histone proteins containing the histone fold motif. These sequences are both eukaryotic and archaeal in origin. Several new histone fold-containing proteins have been identified, including Spt7p, and a few false positives have been removed from the earlier version of HDB. Database contents include compilations of post-translational modifications for each of the core and linker histones, as well as genomic information in the form of map loci for the human histone gene complement, with the genetic loci linked to Online Mendelian Inheritance in Man (OMIM). Conflicts between similar sequence entries from a number of source databases are also documented. Newly added to the HDB are multiple sequence alignments in which predicted functions of histone fold amino acid residues are annotated. The database is freely accessible through the WWW at http://genome.nhgri.nih.gov/histones/

Amino Acid Sequence↗

Fuzzy measures on the Gene Ontology for gene product similarity.

One of the most important objects in bioinformatics is a gene product (protein or RNA). For many gene products, functional information is summarized in a set of Gene Ontology (GO) annotations. For these genes, it is reasonable to include similarity measures based on the terms found in the GO or other taxonomy. In this paper, we introduce several novel measures for computing the similarity of two gene products annotated with GO terms. The fuzzy measure similarity (FMS) has the advantage that it takes into consideration the context of both complete sets of annotation terms when computing the similarity between two gene products. When the two gene products are not annotated by common taxonomy terms, we propose a method that avoids a zero similarity result. To account for the variations in the annotation reliability, we propose a similarity measure based on the Choquet integral. These similarity measures provide extra tools for the biologist in search of functional information for gene products. The initial testing on a group of 194 sequences representing three proteins families shows a higher correlation of the FMS and Choquet similarities to the BLAST sequence similarities than the traditional similarity measures such as pairwise average or pairwise maximum.

Algorithms↗

Enhancer trapping in zebrafish using the Sleeping Beauty transposon.

BACKGROUND: Among functional elements of a metazoan gene, enhancers are particularly difficult to find and annotate. Pioneering experiments in Drosophila have demonstrated the value of enhancer "trapping" using an invertebrate to address this functional genomics problem. RESULTS: We modulated a Sleeping Beauty transposon-based transgenesis cassette to establish an enhancer trapping technique for use in a vertebrate model system, zebrafish Danio rerio. We established 9 lines of zebrafish with distinct tissue- or organ-specific GFP expression patterns from 90 founders that produced GFP-expressing progeny. We have molecularly characterized these lines and show that in each line, a specific GFP expression pattern is due to a single transposition event. Many of the insertions are into introns of zebrafish genes predicted in the current genome assembly. We have identified both previously characterized as well as novel expression patterns from this screen. For example, the ET7 line harbors a transposon insertion near the mkp3 locus and expresses GFP in the midbrain-hindbrain boundary, forebrain and the ventricle, matching a subset of the known FGF8-dependent mkp3 expression domain. The ET2 line, in contrast, expresses GFP specifically in caudal primary motoneurons due to an insertion into the poly(ADP-ribose) glycohydrolase (PARG) locus. This surprising expression pattern was confirmed using in situ hybridization techniques for the endogenous PARG mRNA, indicating the enhancer trap has replicated this unexpected and highly localized PARG expression with good fidelity. Finally, we show that it is possible to excise a Sleeping Beauty transposon from a genomic location in the zebrafish germline. CONCLUSIONS: This genomics tool offers the opportunity for large-scale biological approaches combining both expression and genomic-level sequence analysis using as a template an entire vertebrate genome.

Animals↗

Database for chicken full-length cDNAs.

The generation of full-length cDNA databases is essential for functional genomics studies as well as for correct annotation of species genomic sequences. Human and mouse full-length cDNA projects have provided the biomedical research community with a large amount of gene information. Recent completion of the chicken genome sequence draft now enables a similar full-length cDNA project to be initiated for this species. In this report, we introduce the development of a chicken full-length cDNA database, which will facilitate future research work in this biological system. In this project, chicken expressed sequence tags (ESTs) were aligned onto human and mouse full-length cDNAs (or open reading frames) on the basis of their similarity. More than 588,000 chicken ESTs were aligned to approximately 170,000 full-length human and mouse templates obtained from the NEDO, RIKEN, and MGC databases. Many of these templates have known biological functions, and their orthologous chicken genes in the EMBL database are also provided in our database, which is available at http://bioinfo.hku.hk/chicken/. We will continue to collect known chicken full-length cDNAs to update the database for public use. The cDNA alignment results presented herein and on our database will be useful for animal science and veterinary researchers wishing to clone and to confirm full-length chicken cDNAs of interest.

Animals↗

DNA replication-timing analysis of human chromosome 22 at high resolution and different developmental states.

Duplication of the genome during the S phase of the cell cycle does not occur simultaneously; rather, different sequences are replicated at different times. The replication timing of specific sequences can change during development; however, the determinants of this dynamic process are poorly understood. To gain insights into the contribution of developmental state, genomic sequence, and transcriptional activity to replication timing, we investigated the timing of DNA replication at high resolution along an entire human chromosome (chromosome 22) in two different cell types. The pattern of replication timing was correlated with respect to annotated genes, gene expression, novel transcribed regions of unknown function, sequence composition, and cytological features. We observed that chromosome 22 contains regions of early- and late-replicating domains of 100 kb to 2 Mb, many (but not all) of which are associated with previously described chromosomal bands. In both cell types, expressed sequences are replicated earlier than nontranscribed regions. However, several highly transcribed regions replicate late. Overall, the DNA replication-timing profiles of the two different cell types are remarkably similar, with only nine regions of difference observed. In one case, this difference reflects the differential expression of an annotated gene that resides in this region. Novel transcribed regions with low coding potential exhibit a strong propensity for early DNA replication. Although the cellular function of such transcripts is poorly understood, our results suggest that their activity is linked to the replication-timing program.

Cell Differentiation↗

The proteome: structure, function and evolution.

This paper reports two studies to model the inter-relationships between protein sequence, structure and function. First, an automated pipeline to provide a structural annotation of proteomes in the major genomes is described. The results are stored in a database at Imperial College, London (3D-GENOMICS) that can be accessed at www.sbg.bio.ic.ac.uk. Analysis of the assignments to structural superfamilies provides evolutionary insights. 3D-GENOMICS is being integrated with related proteome annotation data at University College London and the European Bioinformatics Institute in a project known as e-protein (http://www.e-protein.org/). The second topic is motivated by the developments in structural genomics projects in which the structure of a protein is determined prior to knowledge of its function. We have developed a new approach PHUNCTIONER that uses the gene ontology (GO) classification to supervise the extraction of the sequence signal responsible for protein function from a structure-based sequence alignment. Using GO we can obtain profiles for a range of specificities described in the ontology. In the region of low sequence similarity (around 15%), our method is more accurate than assignment from the closest structural homologue. The method is also able to identify the specific residues associated with the function of the protein family.

Computational Biology↗

rVista for comparative sequence-based discovery of functional transcription factor binding sites.

Identifying transcriptional regulatory elements represents a significant challenge in annotating the genomes of higher vertebrates. We have developed a computational tool, rVista, for high-throughput discovery of cis-regulatory elements that combines clustering of predicted transcription factor binding sites (TFBSs) and the analysis of interspecies sequence conservation to maximize the identification of functional sites. To assess the ability of rVista to discover true positive TFBSs while minimizing the prediction of false positives, we analyzed the distribution of several TFBSs across 1 Mb of the well-annotated cytokine gene cluster (Hs5q31; Mm11). Because a large number of AP-1, NFAT, and GATA-3 sites have been experimentally identified in this interval, we focused our analysis on the distribution of all binding sites specific for these transcription factors. The exploitation of the orthologous human-mouse dataset resulted in the elimination of > 95% of the approximately 58,000 binding sites predicted on analysis of the human sequence alone, whereas it identified 88% of the experimentally verified binding sites in this region.

Animals↗

The complement of enzymatic sets in different species.

We present here a comprehensive analysis of the complement of enzymes in a large variety of species. As enzymes are a relatively conserved group there are several classification systems available that are common to all species and link a protein sequence to an enzymatic function. Enzymes are therefore an ideal functional group to study the relationship between sequence expansion, functional divergence and phenotypic changes. By using information retrieved from the well annotated SWISS-PROT database together with sequence information from a variety of fully sequenced genomes and information from the EC functional scheme we have aimed here to estimate the fraction of enzymes in genomes, to determine the extent of their functional redundancy in different domains of life and to identify functional innovations and lineage specific expansions in the metazoa lineage. We found that prokaryote and eukaryote species differ both in the fraction of enzymes in their genomes and in the pattern of expansion of their enzymatic sets. We observe an increase in functional redundancy accompanying an increase in species complexity. A quantitative assessment was performed in order to determine the degree of functional redundancy in different species. Finally, we report a massive expansion in the number of mammalian enzymes involved in signalling and degradation.

Animals↗

Serial analysis of gene expression in sugarcane (Saccharum spp.) leaves revealed alternative C4 metabolism and putative antisense transcripts.

Sugarcane (Saccharum spp.) is a highly efficient biomass and sugar producing crop. Leaf reactions have been considered as potential rate-limiting step for sucrose accumulation in sugarcane stalks. To characterize the sugarcane leaf transcriptome, field-grown mature leaves from cultivar "SP80-3280" were analyzed using Serial Analysis of Gene Expression (SAGE). From 480 sequenced clones, 9,482 valid tags were extracted, with 5,227 unique sequences, from which 3,659 (70%) matched at least a sugarcane assembled sequence (SAS) with putative function; while 872 tags (16.7%) matched SAS with unknown function; 523 (10%) matched SAS without a putative annotation; and only 173 (3.3%) did not match any sugarcane ESTs. Based on gene ontology (GO), photosystem (PS) I reaction center was identified as the most frequent gene product location, followed by the remaining sites of PS I, PS II and thylakoid complexes. For metabolic processes, photosynthesis light harvesting complexes; carbon fixation; and chlorophyll biosynthesis were the most enriched GO-terms. Considering the alternative photosynthetic C(4) cycles, tag frequencies related to phosphoenolpyruvate carboxykinase (PEPCK) and aspartate aminotransferase compared to those for NADP(+)-malic enzyme (NADP-ME) and NADP-malate dehydrogenase, suggested that PEPCK-type decarboxylation appeared to predominate over NADP-ME in mature leaves, although both may occur, opposite to currently assumed in sugarcane. From the unique tag set, 894 tags (17.1%) were assigned as potentially derived from antisense transcripts, while 73 tags (1.4%) were assigned to more than one SAS, suggesting the occurrence of alternative processing. The occurrence of antisense was validated by quantitative reverse transcription amplification. Sugarcane leaf transcriptome provided new insights for functional studies associated with sucrose synthesis and accumulation.

Carbohydrate Metabolism↗

PathFinder: reconstruction and dynamic visualization of metabolic pathways.

MOTIVATION: Beyond methods for a gene-wise annotation and analysis of sequenced genomes new automated methods for functional analysis on a higher level are needed. The identification of realized metabolic pathways provides valuable information on gene expression and regulation. Detection of incomplete pathways helps to improve a constantly evolving genome annotation or discover alternative biochemical pathways. To utilize automated genome analysis on the level of metabolic pathways new methods for the dynamic representation and visualization of pathways are needed. RESULTS: PathFinder is a tool for the dynamic visualization of metabolic pathways based on annotation data. Pathways are represented as directed acyclic graphs, graph layout algorithms accomplish the dynamic drawing and visualization of the metabolic maps. A more detailed analysis of the input data on the level of biochemical pathways helps to identify genes and detect improper parts of annotations. As an Relational Database Management System (RDBMS) based internet application PathFinder reads a list of EC-numbers or a given annotation in EMBL- or Genbank-format and dynamically generates pathway graphs.

Bacillus subtilis↗

Structural domains, protein modules, and sequence similarities enrich our understanding of the Shewanella oneidensis MR-1 proteome.

The protein coding sequences of S. oneidensis MR-1 were analyzed, and new annotations were given to 491 gene products, 306 of which were previously of unknown function. New information was mainly brought in from structural domain predictions for S. oneidensis proteins of the SUPERFAM database (http://supfam.mrc-lmb.cam.ac.uk/SUPERFAMILY/) and newly identified and experimentally verified functions of homologous proteins. Proteins encoded by fused genes were identified and separated into modules, protein units of at least 83 aa with independent functions and distinct evolutionary histories. A reannotation of the fused gene products was done to assign functions to the appropriate module within the protein. Groups of sequence-similar proteins of S. oneidensis were assembled. The fused gene products were represented by their modular entities for the grouping process. The protein groups were analyzed for their size and functions, and they were used to indicate activities that are of importance to the environmental adaptation of this organism. Making use of several approaches not commonly used in annotation, we have been able to enrich our understanding of the functions encoded by the S. oneidensis genome.

Bacterial Proteins↗

Recognizing complex, asymmetric functional sites in protein structures using a Bayesian scoring function.

The increase in known three-dimensional protein structures enables us to build statistical profiles of important functional sites in protein molecules. These profiles can then be used to recognize sites in large-scale automated annotations of new protein structures. We report an improved FEATURE system which recognizes functional sites in protein structures. FEATURE defines multi-level physico-chemical properties and recognizes sites based on the spatial distribution of these properties in the sites' microenvironments. It uses a Bayesian scoring function to compare a query region with the statistical profile built from known examples of sites and control nonsites. We have previously shown that FEATURE can accurately recognize calcium-binding sites and have reported interesting results scanning for calcium-binding sites in the entire Protein Data Bank. Here we report the ability of the improved FEATURE to characterize and recognize geometrically complex and asymmetric sites such as ATP-binding sites and disulfide bond-forming sites. FEATURE does not rely on conserved residues or conserved residue geometry of the sites. We also demonstrate that, in the absence of a statistical profile of the sites, FEATURE can use an artificially constructed profile based on a priori knowledge to recognize the sites in new structures, using redoxin active sites as an example.

Adenosine Triphosphate↗