Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

HyPaLib: a database of RNAs and RNA structural elements defined by hybrid patterns.

The database, called HyPaLib (for Hybrid Pattern Library), contains annotated structural elements characteristic for certain classes of structural and/or functional RNAs. These elements are described in a language specifically designed for this purpose. The language allows convenient specification of hybrid patterns, i.e. motifs consisting of sequence features and structural elements together with sequence similarity and thermodynamic constraints. We are currently developing software tools that allow a user to search sequence databases for any pattern in HyPaLib, thus providing functionality which is similar to PROSITE, but dedicated to the more complex patterns in RNA sequences. HyPaLib is available at http://bibiserv. techfak.uni-bielefeld.de/HyPa/.

Algorithms↗

Transcriptome analysis of Escherichia coli using high-density oligonucleotide probe arrays.

Microarrays traditionally have been used to analyze the expression behavior of large numbers of coding transcripts. Here we present a comprehensive approach for high-throughput transcript discovery in Escherichia coli focused mainly on intergenic regions which, together with analysis of coding transcripts, provides us with a more complete insight into the organism's transcriptome. Using a whole genome array, we detected expression for 4052 coding transcripts and identified 1102 additional transcripts in the intergenic regions of the E.coli genome. Further classification reveals 317 novel transcripts with unknown function. Our results show that, despite sophisticated approaches to genome annotation, many cellular transcripts remain unidentified. Through the experimental identification of all RNAs expressed under a specific condition, we gain a more thorough understanding of all cellular processes.

3' Untranslated Regions↗

Organelle DB: an updated resource of eukaryotic protein localization and function.

Organelle DB (http://organelledb.lsi.umich.edu) is a web-accessible relational database presenting a supplemented catalog of organelle-localized proteins and major protein complexes. Since its release in 2004, Organelle DB has grown by 20% to encompass over 30,000 proteins from 138 eukaryotic organisms. Each protein in Organelle DB is presented with its subcellular localization, primary sequence and a detailed description of its function, as available. All records in Organelle DB have been annotated using controlled vocabulary from the Gene Ontology consortium. Protein localization data are inherently visual, and Organelle DB is a significant repository of biological images, housing 1500 micrographs of yeast cells carrying stained proteins. Furthermore, we report here the development of Organelle View, an extension of Organelle DB for the interactive visualization of organelles and subcellular structures in the budding yeast Saccharomyces cerevisiae. Organelle View offers a dimensional representation of a yeast cell; users can search Organelle View for proteins of interest, and the organelles housing these proteins will be highlighted in the cell image. Among other applications, Organelle View may serve as an educational aid engaging introductory biology students through a visually 'fun' interface. Organelle View can be accessed from the Organelle DB home page or directly at http://organelleview.lsi.umich.edu.

Animals↗

SNP@Ethnos: a database of ethnically variant single-nucleotide polymorphisms.

Inherited genetic variation plays a critical but largely uncharacterized role in human differentiation. The completion of the International HapMap Project makes it possible to identify loci that may cause human differentiation. We have devised an approach to find such ethnically variant single-nucleotide polymorphisms (ESNPs) from the genotype profile of the populations included in the International HapMap database. We selected ESNPs using the nearest shrunken centroid method (NSCM), and performed multiple tests for genetic heterogeneity and frequency spectrum on genes having ESNPs. The function and disease association of the selected SNPs were also annotated. This resulted in the identification of 100 736 SNPs that appeared uniquely in each ethnic group. Of these SNPs, 1009 were within disease-associated genes, and 85 were predicted as damaging using the Sorting Intolerant From Tolerant system. This study resulted in the creation of the SNP@Ethnos database, which is designed to make this type of detailed genetic variation approach available to a wider range of researchers. SNP@Ethnos is a public database of ESNPs with annotation information that currently contains 100 736 ESNPs from 10 138 genes, and can be accessed at http://variome.net and http://bioportal.net/ or directly at http://bioportal.kobic.re.kr/SNPatETHNIC/.

Chromosome Mapping↗

Identification of a novel gene, URE2, that functionally complements a urease-negative clinical strain of Cryptococcus neoformans.

A urease-negative serotype A strain of Cryptococcus neoformans (B-4587) was isolated from the cerebrospinal fluid of an immunocompetent patient with a central nervous system infection. The URE1 gene encoding urease failed to complement the mutant phenotype. Urease-positive clones of B-4587 obtained by complementing with a genomic library of strain H99 harboured an episomal plasmid containing DNA inserts with homology to the sudA gene of Aspergillus nidulans. The gene harboured by these plasmids was named URE2 since it enabled the transformants to grow on media containing urea as the sole nitrogen source while the transformants with an empty vector failed to grow. Transformation of strain B-4587 with a plasmid construct containing a truncated version of the URE2 gene failed to complement the urease-negative phenotype. Disruption of the native URE2 gene in a wild-type serotype A strain H99 and a serotype D strain LP1 of C. neoformans resulted in the inability of the strains to grow on media containing urea as the sole nitrogen source, suggesting that the URE2 gene product is involved in the utilization of urea by the organism. Virulence in mice of the urease-negative isolate B-4587, the urease-positive transformants containing the wild-type copy of the URE2 gene, and the urease-negative vector-only transformants was comparable to that of the H99 strain of C. neoformans regardless of the infection route. Virulence of the URE2 disruption stain of H99 was slightly reduced compared to the wild-type strain in the intravenous model but was significantly attenuated in the inhalation model. These results indicate that the importance of urease activity in pathogenicity varies depending on the strains of C. neoformans used and/or the route of infection. Furthermore, this study shows that complementation cloning can serve as a useful tool to functionally identify genes such as URE2 that have otherwise been annotated as hypothetical proteins in genomic databases.

Animals↗

Identification, analysis, and utilization of conserved ortholog set markers for comparative genomics in higher plants.

We have screened a large tomato EST database against the Arabidopsis genomic sequence and report here the identification of a set of 1025 genes (referred to as a conserved ortholog set, or COS markers) that are single or low copy in both genomes (as determined by computational screens and DNA gel blot hybridization) and that have remained relatively stable in sequence since the early radiation of dicotyledonous plants. These genes were annotated, and a large portion could be assigned to putative functional categories associated with basic metabolic processes, such as energy-generating processes and the biosynthesis and degradation of cellular building blocks. We further demonstrate, through computational screens (e.g., against a Medicago truncatula database) and direct hybridization on genomic DNA of diverse plant species, that these COS markers also are conserved in the genomes of other plant families. Finally, we show that this gene set can be used for comparative mapping studies between highly divergent genomes such as those of tomato and Arabidopsis. This set of COS markers, identified computationally and experimentally, may further studies on comparative genomes and phylogenetics and elucidate the nature of genes conserved throughout plant evolution.

Amino Acid Sequence↗

Microarray and genetic analysis of electron transfer to electrodes in Geobacter sulfurreducens.

Whole-genome analysis of gene expression in Geobacter sulfurreducens revealed 474 genes with transcript levels that were significantly different during growth with an electrode as the sole electron acceptor versus growth on Fe(III) citrate. The greatest response was a more than 19-fold increase in transcript levels for omcS, which encodes an outer-membrane cytochrome previously shown to be required for Fe(III) oxide reduction. Quantitative reverse transcription polymerase chain reaction and Northern analyses confirmed the higher levels of omcS transcripts, which increased as power production increased. Deletion of omcS inhibited current production that was restored when omcS was expressed in trans. Transcript expression and genetic analysis suggested that OmcE, another outer-membrane cytochrome, is also involved in electron transfer to electrodes. Surprisingly, genes for other proteins known to be important in Fe(III) reduction such as the outer-membrane c-type cytochrome, OmcB, and the electrically conductive pilin "nanowires" did not have higher transcript levels on electrodes, and deletion of the relevant genes did not inhibit power production. Changes in the transcriptome suggested that cells growing on electrodes were subjected to less oxidative stress than cells growing on Fe(III) citrate and that a number of genes annotated as encoding metal efflux proteins or proteins of unknown function may be important for growth on electrodes. These results demonstrate for the first time that it is possible to evaluate gene expression, and hence the metabolic state, of microorganisms growing on electrodes on a genome-wide basis and suggest that OmcS, and to a lesser extent OmcE, are important in electron transfer to electrodes. This has important implications for the design of electrode materials and the genetic engineering of microorganisms to improve the function of microbial fuel cells.

Bacterial Outer Membrane Proteins↗

Plant metabolomics: from holistic hope, to hype, to hot topic.

In a short time, plant metabolomics has gone from being just an ambitious concept to being a rapidly growing, valuable technology applied in the stride to gain a more global picture of the molecular organization of multicellular organisms. The combination of improved analytical capabilities with newly designed, dedicated statistical, bioinformatics and data mining strategies, is beginning to broaden the horizons of our understanding of how plants are organized and how metabolism is both controlled but highly flexible. Metabolomics is predicted to play a significant, if not indispensable role in bridging the phenotype-genotype gap and thus in assisting us in our desire for full genome sequence annotation as part of the quest to link gene to function. Plants are a fabulously rich source of diverse functional biochemicals and metabolomics is also already proving valuable in an applied context. By creating unique opportunities for us to interrogate plant systems and characterize their biochemical composition, metabolomics will greatly assist in identifying and defining much of the still unexploited biodiversity available today.

Biotechnology↗

A comparative proteomics resource: proteins of Arabidopsis thaliana.

Using an integrative genome annotation pipeline (iGAP) for proteome-wide protein structure and functional domain assignment, we analyzed all the proteins of Arabidopsis thaliana. Three-dimensional structures at the level of the domain are assigned by fold recognition and threading based on a novel fold library that extends common domain classifications. iGAP is being applied to proteins from all available proteomes as part of a comparative proteomics resource. The database is accessible from the web.

Arabidopsis↗

Transcriptional slippage in bacteria: distribution in sequenced genomes and utilization in IS element gene expression.

BACKGROUND: Transcription slippage occurs on certain patterns of repeat mononucleotides, resulting in synthesis of a heterogeneous population of mRNAs. Individual mRNA molecules within this population differ in the number of nucleotides they contain that are not specified by the template. When transcriptional slippage occurs in a coding sequence, translation of the resulting mRNAs yields more than one protein product. Except where the products of the resulting mRNAs have distinct functions, transcription slippage occurring in a coding region is expected to be disadvantageous. This probably leads to selection against most slippage-prone sequences in coding regions. RESULTS: To find a length at which such selection is evident, we analyzed the distribution of repetitive runs of A and T of different lengths in 108 bacterial genomes. This length varies significantly among different bacteria, but in a large proportion of available genomes corresponds to nine nucleotides. Comparative sequence analysis of these genomes was used to identify occurrences of 9A and 9T transcriptional slippage-prone sequences used for gene expression. CONCLUSIONS: IS element genes are the largest group found to exploit this phenomenon. A number of genes with disrupted open reading frames (ORFs) have slippage-prone sequences at which transcriptional slippage would result in uninterrupted ORF restoration at the mRNA level. The ability of such genes to encode functional full-length protein products brings into question their annotation as pseudogenes and in these cases is pertinent to the significance of the term 'authentic frameshift' frequently assigned to such genes.

Adenosine↗

[Malaria control in the post-genomic era].

With the publication of seminal articles on the genomic sequence of Plasmodium falciparum and P. yoelii yoelii last autumn following the articles on genome of Anopheles gambiae, man, and mice, malariology stepped indubitably into a new era. A major finding of genome annotation is that nearly 60% of Plasmodium sp. have no functional attribution and that a high proportion of P. falciparum sequences have no counterpart in P. yoelii. In the light of these findings it will now be possible to explore the specificities of the Plasmodium genus and particularities of each species--and ultimately the relevance of animal malaria models. Numerous avenues of research have been opened not only for drug discovery, vaccine development, dissecting the cascade of events contributing to pathology or protection, but also for population biology studies and analysis of genetic exchanges within field parasite populations. The field has been radically changed, and coordinated efforts are needed to implement the powerful post genomic technologies. There is an urgent need to exploit this huge mass of new data for the development of new control methods. This development requires a rational approach of functional genomics and integrative biology. The purpose of this article is to briefly summarize current data and discuss experimental approaches and perspectives.

Animals↗

Rates of divergence in gene expression profiles of primates, mice, and flies: stabilizing selection and variability among functional categories.

The extent to which natural selection shapes phenotypic variation has long been a matter of debate among those studying organic evolution. We studied the patterns of gene expression polymorphism and divergence in several datasets that ranged from comparisons between two very closely related laboratory strains of mice to comparisons across a considerably longer time scale, such as between humans and chimpanzees, two species of mice, and two species of Drosophila. The results were analyzed and interpreted in view of neutral models of phenotypic evolution. Our analyses used a number of metrics to show that most mRNA levels are evolutionary stable, changing little across the range of taxonomic distances compared. This implies that, overall, widespread stabilizing selection on transcription levels has prevented greater evolutionary changes in mRNA levels. Nevertheless, the range of rates of divergence is large with highly significant differences in the rate and patterns of transcription divergence across functional classes defined on the basis of the gene ontology annotation (primates and mice datasets) or on the basis of the pattern of sex-biased gene expression (Drosophila). Moreover, rates of divergence of sex-biased genes in the contrast between Drosophila species show a distinct pattern from that observed in the contrast between populations of D. melanogaster. Hence, we discuss the time scale of the changes observed and its consequences for the relationship between variation in gene expression within and between species. Finally, we argue that differences in mRNA levels of the magnitudes observed herein could be explained by a remarkably small number of generations of directional selection.

Animals↗

Functional reclassification of the putative cinnamyl alcohol dehydrogenase multigene family in Arabidopsis.

Of 17 genes annotated in the Arabidopsis genome database as cinnamyl alcohol dehydrogenase (CAD) homologues, an in silico analysis revealed that 8 genes were misannotated. Of the remaining nine, six were catalytically competent for NADPH-dependent reduction of p-coumaryl, caffeyl, coniferyl, 5-hydroxyconiferyl, and sinapyl aldehydes, whereas three displayed very low activity and only at very high substrate concentrations. Of the nine putative CADs, two (AtCAD5 and AtCAD4) had the highest activity and homology (approximately 83% similarity) relative to bona fide CADs from other species. AtCAD5 used all five substrates effectively, whereas AtCAD4 (of lower overall catalytic capacity) poorly used sinapyl aldehyde; the corresponding 270-fold decrease in k(enz) resulted from higher K(m) and lower k(cat) values, respectively. No CAD homologue displayed a specific requirement for sinapyl aldehyde, which was in direct contrast with unfounded claims for a so-called sinapyl alcohol dehydrogenase in angiosperms. AtCAD2, 3, as well as AtCAD7 and 8 (highest homology to sinapyl alcohol dehydrogenase) were catalytically less active overall by at least an order of magnitude, due to increased K(m) and lower k(cat) values. Accordingly, alternative and/or bifunctional metabolic roles of these proteins in plant defense cannot be ruled out. Comprehensive analyses of lignified tissues of various Arabidopsis knockout mutants (for AtCAD5, 6, and 9) at different stages of growth/development indicated the presence of functionally redundant CAD metabolic networks. Moreover, disruption of AtCAD5 expression had only a small effect on either overall lignin amounts deposited, or on syringyl-guaiacyl compositions, despite being the most catalytically active form in vitro.

Alcohol Dehydrogenase↗

Optimal cDNA microarray design using expressed sequence tags for organisms with limited genomic information.

BACKGROUND: Expression microarrays are increasingly used to characterize environmental responses and host-parasite interactions for many different organisms. Probe selection for cDNA microarrays using expressed sequence tags (ESTs) is challenging due to high sequence redundancy and potential cross-hybridization between paralogous genes. In organisms with limited genomic information, like marine organisms, this challenge is even greater due to annotation uncertainty. No general tool is available for cDNA microarray probe selection for these organisms. Therefore, the goal of the design procedure described here is to select a subset of ESTs that will minimize sequence redundancy and characterize potential cross-hybridization while providing functionally representative probes. RESULTS: Sequence similarity between ESTs, quantified by the E-value of pair-wise alignment, was used as a surrogate for expected hybridization between corresponding sequences. Using this value as a measure of dissimilarity, sequence redundancy reduction was performed by hierarchical cluster analyses. The choice of how many microarray probes to retain was made based on an index developed for this research: a sequence diversity index (SDI) within a sequence diversity plot (SDP). This index tracked the decreasing within-cluster sequence diversity as the number of clusters increased. For a given stage in the agglomeration procedure, the EST having the highest similarity to all the other sequences within each cluster, the centroid EST, was selected as a microarray probe. A small dataset of ESTs from Atlantic white shrimp (Litopenaeus setiferus) was used to test this algorithm so that the detailed results could be examined. The functional representative level of the selected probes was quantified using Gene Ontology (GO) annotations. CONCLUSIONS: For organisms with limited genomic information, combining hierarchical clustering methods to analyze ESTs can yield an optimal cDNA microarray design. If biomarker discovery is the goal of the microarray experiments, the average linkage method is more effective, while single linkage is more suitable if identification of physiological mechanisms is more of interest. This general design procedure is not limited to designing single-species cDNA microarrays for marine organisms, and it can equally be applied to multiple-species microarrays of any organisms with limited genomic information.

Animals↗

Exploiting conserved structure for faster annotation of non-coding RNAs without loss of accuracy.

MOTIVATION: Non-coding RNAs (ncRNAs)-functional RNA molecules not coding for proteins-are grouped into hundreds of families of homologs. To find new members of an ncRNA gene family in a large genome database, covariance models (CMs) are a useful statistical tool, as they use both sequence and RNA secondary structure information. Unfortunately, CM searches are slow. Previously, we introduced 'rigorous filters', which provably sacrifice none of CMs' accuracy, although often scanning much faster. A rigorous filter, using a profile hidden Markov model (HMM), is built based on the CM, and filters the genome database, eliminating sequences that provably could not be annotated as homologs. The CM is run only on the remainder. Some biologically important ncRNA families could not be scanned efficiently with this technique, largely due to the significance of conserved secondary structure relative to primary sequence in identifying these families. Current heuristic filters are also expected to perform poorly on such families. RESULTS: By augmenting profile HMMs with limited secondary structure information, we obtain rigorous filters that accelerate CM searches for virtually all known ncRNA families from the Rfam Database and tRNA models in tRNAscan-SE. These filters scan an 8 gigabase database in weeks instead of years, and uncover homologs missed by heuristic techniques to speed CM searches. AVAILABILITY: Software in development; contact the authors.

Algorithms↗

Experimental validation of novel genes predicted in the un-annotated regions of the Arabidopsis genome.

BACKGROUND: Several lines of evidence support the existence of novel genes and other transcribed units which have not yet been annotated in the Arabidopsis genome. Two gene prediction programs which make use of comparative genomic analysis, Twinscan and EuGene, have recently been deployed on the Arabidopsis genome. The ability of these programs to make use of sequence data from other species has allowed both Twinscan and EuGene to predict over 1000 genes that are intergenic with respect to the most recent annotation release. A high throughput RACE pipeline was utilized in an attempt to verify the structure and expression of these novel genes. RESULTS: 1,071 un-annotated loci were targeted by RACE, and full length sequence coverage was obtained for 35% of the targeted genes. We have verified the structure and expression of 378 genes that were not present within the most recent release of the Arabidopsis genome annotation. These 378 genes represent a structurally diverse set of transcripts and encode a functionally diverse set of proteins. CONCLUSION: We have investigated the accuracy of the Twinscan and EuGene gene prediction programs and found them to be reliable predictors of gene structure in Arabidopsis. Several hundred previously un-annotated genes were validated by this work. Based upon this information derived from these efforts it is likely that the Arabidopsis genome annotation continues to overlook several hundred protein coding genes.

Arabidopsis↗

Non-lexical approaches to identifying associative relations in the gene ontology.

The Gene Ontology (GO) is a controlled vocabulary widely used for the annotation of gene products. GO is organized in three hierarchies for molecular functions, cellular components, and biological processes but no relations are provided among terms across hierarchies. The objective of this study is to investigate three non-lexical approaches to identifying such associative relations in GO and compare them among themselves and to lexical approaches. The three approaches are: computing similarity in a vector space model, statistical analysis of co-occurrence of GO terms in annotation databases, and association rule mining. Five annotation databases (FlyBase, the Human subset of GOA, MGI, SGD, and WormBase) are used in this study. A total of 7,665 associations were identified by at least one of the three non-lexical approaches. Of these, 12% were identified by more than one approach. While there are almost 6,000 lexical relations among GO terms, only 203 associations were identified by both non-lexical and lexical approaches. The associations identified in this study could serve as the starting point for adding associative relations across hierarchies to GO, but would require manual curation. The application to quality assurance of annotation databases is also discussed.

Alzheimer Disease↗

Representation of roles in biomedical ontologies: a case study in functional genomics.

OBJECTIVE: Representing roles, i.e. functions of proteins, sequences and structures, is the cornerstone of knowledge representation in functional genomics. The objective of this study is to investigate representation of roles as functional categories or associative relations. We focus on GeneOntology (GO) and the UMLS and take examples from iron metabolism. METHODS: The terms corresponding to the main proteins involved in iron metabolism were mapped to GO (including the annotations) and the UMLS. The representation of their biological roles was then analyzed. RESULTS: Functional aspects are represented in both GO and the UMLS. However, the granularity may not be appropriate. DISCUSSION: Advantages and limits of functional categories and associative relations are discussed.

Genes↗