Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,801 records · Page 100Linked to original sources

Adult mouse brain gene expression patterns bear an embryologic imprint.

The current model to explain the organization of the mammalian nervous system is based on studies of anatomy, embryology, and evolution. To further investigate the molecular organization of the adult mammalian brain, we have built a gene expression-based brain map. We measured gene expression patterns for 24 neural tissues covering the mouse central nervous system and found, surprisingly, that the adult brain bears a transcriptional "imprint" consistent with both embryological origins and classic evolutionary relationships. Embryonic cellular position along the anterior-posterior axis of the neural tube was shown to be closely associated with, and possibly a determinant of, the gene expression patterns in adult structures. We also observed a significant number of embryonic patterning and homeobox genes with region-specific expression in the adult nervous system. The relationships between global expression patterns for different anatomical regions and the nature of the observed region-specific genes suggest that the adult brain retains a degree of overall gene expression established during embryogenesis that is important for regional specificity and the functional relationships between regions in the adult. The complete collection of extensively annotated gene expression data along with data mining and visualization tools have been made available on a publicly accessible web site (www.barlow-lockhart-brainmapnimhgrant.org).

Algorithms↗

A simple algorithm to infer gene duplication and speciation events on a gene tree.

MOTIVATION: When analyzing protein sequences using sequence similarity searches, orthologous sequences (that diverged by speciation) are more reliable predictors of a new protein's function than paralogous sequences (that diverged by gene duplication), because duplication enables functional diversification. The utility of phylogenetic information in high-throughput genome annotation ('phylogenomics') is widely recognized, but existing approaches are either manual or indirect (e.g. not based on phylogenetic trees). Our goal is to automate phylogenomics using explicit phylogenetic inference. A necessary component is an algorithm to infer speciation and duplication events in a given gene tree. RESULTS: We give an algorithm to infer speciation and duplication events on a gene tree by comparison to a trusted species tree. This algorithm has a worst-case running time of O(n(2)) which is inferior to two previous algorithms that are approximately O(n) for a gene tree of sequences. However, our algorithm is extremely simple, and its asymptotic worst case behavior is only realized on pathological data sets. We show empirically, using 1750 gene trees constructed from the Pfam protein family database, that it appears to be a practical (and often superior) algorithm for analyzing real gene trees. AVAILABILITY: http://www.genetics.wustl.edu/eddy/forester.

Algorithms↗

Discovering statistically significant biclusters in gene expression data.

In gene expression data, a bicluster is a subset of the genes exhibiting consistent patterns over a subset of the conditions. We propose a new method to detect significant biclusters in large expression datasets. Our approach is graph theoretic coupled with statistical modelling of the data. Under plausible assumptions, our algorithm is polynomial and is guaranteed to find the most significant biclusters. We tested our method on a collection of yeast expression profiles and on a human cancer dataset. Cross validation results show high specificity in assigning function to genes based on their biclusters, and we are able to annotate in this way 196 uncharacterized yeast genes. We also demonstrate how the biclusters lead to detecting new concrete biological associations. In cancer data we are able to detect and relate finer tissue types than was previously possible. We also show that the method outperforms the biclustering algorithm of Cheng and Church (2000).

Algorithms↗

Genome Properties: a system for the investigation of prokaryotic genetic content for microbiology, genome annotation and comparative genomics.

MOTIVATION: The presence or absence of metabolic pathways and structures provide a context that makes protein annotation far more reliable. Compiling such information across microbial genomes improves the functional classification of proteins and provides a valuable resource for comparative genomics. RESULTS: We have created a Genome Properties system to present key aspects of prokaryotic biology using standardized computational methods and controlled vocabularies. Properties reflect gene content, phenotype, phylogeny and computational analyses. The results of searches using hidden Markov models allow many properties to be deduced automatically, especially for families of proteins (equivalogs) conserved in function since their last common ancestor. Additional properties are derived from curation, published reports and other forms of evidence. Genome Properties system was applied to 156 complete prokaryotic genomes, and is easily mined to find differences between species, correlations between metabolic features and families of uncharacterized proteins, or relationships among properties. AVAILABILITY: Genome Properties can be found at http://www.tigr.org/Genome_Properties SUPPLEMENTARY INFORMATION: http://www.tigr.org/tigr-scripts/CMR2/genome_properties_references.spl.

Chromosome Mapping↗

MIPS Arabidopsis thaliana Database (MAtDB): an integrated biological knowledge resource for plant genomics.

Arabidopsis thaliana is the most widely studied model plant. Functional genomics is intensively underway in many laboratories worldwide. Beyond the basic annotation of the primary sequence data, the annotated genetic elements of Arabidopsis must be linked to diverse biological data and higher order information such as metabolic or regulatory pathways. The MIPS Arabidopsis thaliana database MAtDB aims to provide a comprehensive resource for Arabidopsis as a genome model that serves as a primary reference for research in plants and is suitable for transfer of knowledge to other plants, especially crops. The genome sequence as a common backbone serves as a scaffold for the integration of data, while, in a complementary effort, these data are enhanced through the application of state-of-the-art bioinformatics tools. This information is visualized on a genome-wide and a gene-by-gene basis with access both for web users and applications. This report updates the information given in a previous report and provides an outlook on further developments. The MAtDB web interface can be accessed at http://mips.gsf.de/proj/thal/db.

Arabidopsis↗

PHI-base: a new database for pathogen host interactions.

To utilize effectively the growing number of verified genes that mediate an organism's ability to cause disease and/or to trigger host responses, we have developed PHI-base. This is a web-accessible database that currently catalogs 405 experimentally verified pathogenicity, virulence and effector genes from 54 fungal and Oomycete pathogens, of which 176 are from animal pathogens, 227 from plant pathogens and 3 from pathogens with a fungal host. PHI-base is the first on-line resource devoted to the identification and presentation of information on fungal and Oomycete pathogenicity genes and their host interactions. As such, PHI-base is a valuable resource for the discovery of candidate targets in medically and agronomically important fungal and Oomycete pathogens for intervention with synthetic chemistries and natural products. Each entry in PHI-base is curated by domain experts and supported by strong experimental evidence (gene/transcript disruption experiments) as well as literature references in which the experiments are described. Each gene in PHI-base is presented with its nucleotide and deduced amino acid sequence as well as a detailed description of the predicted protein's function during the host infection process. To facilitate data interoperability, we have annotated genes using controlled vocabularies (Gene Ontology terms, Enzyme Commission Numbers and so on), and provide links to other external data sources (e.g. NCBI taxonomy and EMBL). We welcome new data for inclusion in PHI-base, which is freely accessed at www4.rothamsted.bbsrc.ac.uk/phibase/.

Algal Proteins↗

Evolution of transcription factor binding sites in Mammalian gene regulatory regions: conservation and turnover.

Comparisons between human and rodent DNA sequences are widely used for the identification of regulatory regions (phylogenetic footprinting), and the importance of such intergenomic comparisons for promoter annotation is expanding. The efficacy of such comparisons for the identification of functional regulatory elements hinges on the evolutionary dynamics of promoter sequences. Although it is widely appreciated that conservation of sequence motifs may provide a suggestion of function, it is not known as to what proportion of the functional binding sites in humans is conserved in distant species. In this report, we present an analysis of the evolutionary dynamics of transcription factor binding sites whose function had been experimentally verified in promoters of 51 human genes and compare their sequence to homologous sequences in other primate species and rodents. Our results show that there is extensive divergence within the nucleotide sequence of transcription factor binding sites. Using direct experimental data from functional studies in both human and rodents for 20 of the regulatory regions, we estimate that 32%-40% of the human functional sites are not functional in rodents. This is evidence that there is widespread turnover of transcription factor binding sites. These results have important implications for the efficacy of phylogenetic footprinting and the interpretation of the pattern of evolution in regulatory sequences.

Animals↗

Genomic survey of cAMP and cGMP signalling components in the cyanobacterium Synechocystis PCC 6803.

Cyanobacteria modulate intracellular levels of cAMP and cGMP in response to environmental conditions (light, nutrients and pH). In an attempt to identify components of the cAMP and cGMP signalling pathways in Synechocystis PCC 6803, the authors screened its complete genome sequence by using bioinformatic tools and data from sequence-function studies performed on both eukaryotic and prokaryotic cAMP/cGMP-dependent proteins. Sll1624 and Slr2100 were tentatively assigned as being two putative cyclic nucleotide phosphodiesterases. Five proteins were identified as having all the determinants required to be cyclic nucleotide receptors, two of them being probably more specific for cGMP (an element of two-component regulatory systems - Slr2104 - and a putative cyclic-nucleotide-gated cation channel - Slr1575), the three others being probably more specific for cAMP: (i) a protein of unidentified function (Slr0842); (ii) a putative cyclic-nucleotide-modulated permease (Slr0593), previously annotated as a kinase A regulatory subunit; and (iii) a putative transcription factor (CRP-SYN: =Sll1371), which possesses cAMP- and DNA-binding determinants homologous to those of the cAMP receptor protein of Escherichia coli (CRP-EC:). This homology, together with the presence in Synechocystis of CRP-EC:-like binding sites upstream of crp, cya1, slr1575, and several genes encoding enzymes involved in transport and metabolism, strongly suggests that CRP-SYN: is a global regulator.

3',5'-Cyclic-AMP Phosphodiesterases↗

MAGPIE/EGRET annotation of the 2.9-Mb Drosophila melanogaster Adh region.

Our challenge in annotating the 2.91-Mb Adh region of the Drosophila melanogaster genome was to identify genetic and genomic features automatically, completely, and precisely within a 6-week period. To do so, we augmented the MAGPIE microbial genome annotation system to handle eukaryotic genomic sequence data. The new configuration required the integration of eukaryotic gene-finding tools and DNA repeat tools into the automatic data collection module. It also required us to define in MAGPIE new strategies to combine data about eukaryotic exon predictions with functional data to refine the exon predictions. At the heart of the resulting new eukaryotic genome annotation system is a reverse comparison of public protein and complementary DNA sequences against the input genome to identify missing exons and to refine exon boundaries. The software modules that add eukaryotic genome annotation capability to MAGPIE are available as EGRET (Eukaryotic Genome Rapid Evaluation Tool).

Alcohol Dehydrogenase↗

Sequence analysis of the 144-kilobase accessory plasmid pSmeSM11a, isolated from a dominant Sinorhizobium meliloti strain identified during a long-term field release experiment.

The genome of Sinorhizobium meliloti type strain Rm1021 consists of three replicons: the chromosome and two megaplasmids, pSymA and pSymB. Additionally, many indigenous S. meliloti strains possess one or more smaller plasmids, which represent the accessory genome of this species. Here we describe the complete nucleotide sequence of an accessory plasmid, designated pSmeSM11a, that was isolated from a dominant indigenous S. meliloti subpopulation in the context of a long-term field release experiment with genetically modified S. meliloti strains. Sequence analysis of plasmid pSmeSM11a revealed that it is 144,170 bp long and has a mean G+C content of 59.5 mol%. Annotation of the sequence resulted in a total of 160 coding sequences. Functional predictions could be made for 43% of the genes, whereas 57% of the genes encode hypothetical or unknown gene products. Two plasmid replication modules, one belonging to the repABC replicon family and the other belonging to the plasmid type A replicator region family, were identified. Plasmid pSmeSM11a contains a mobilization (mob) module composed of the type IV secretion system-related genes traG and traA and a putative mobC gene. A large continuous region that is about 42 kb long is very similar to a corresponding region located on S. meliloti Rm1021 megaplasmid pSymA. Single-base-pair deletions in the homologous regions are responsible for frameshifts that result in nonparalogous coding sequences. Plasmid pSmeSM11a carries additional copies of the nodulation genes nodP and nodQ that are responsible for Nod factor sulfation. Furthermore, a tauD gene encoding a putative taurine dioxygenase was identified on pSmeSM11a. An acdS gene located on pSmeSM11a is the first example of such a gene in S. meliloti. The deduced acdS gene product is able to deaminate 1-aminocyclopropane-1-carboxylate and is proposed to be involved in reducing the phytohormone ethylene, thus influencing nodulation events. The presence of numerous insertion sequences suggests that these elements mediated acquisition of accessory plasmid modules.

Agrobacterium tumefaciens↗

XenDB: full length cDNA prediction and cross species mapping in Xenopus laevis.

BACKGROUND: Research using the model system Xenopus laevis has provided critical insights into the mechanisms of early vertebrate development and cell biology. Large scale sequencing efforts have provided an increasingly important resource for researchers. To provide full advantage of the available sequence, we have analyzed 350,468 Xenopus laevis Expressed Sequence Tags (ESTs) both to identify full length protein encoding sequences and to develop a unique database system to support comparative approaches between X. laevis and other model systems. DESCRIPTION: Using a suffix array based clustering approach, we have identified 25,971 clusters and 40,877 singleton sequences. Generation of a consensus sequence for each cluster resulted in 31,353 tentative contig and 4,801 singleton sequences. Using both BLASTX and FASTY comparison to five model organisms and the NR protein database, more than 15,000 sequences are predicted to encode full length proteins and these have been matched to publicly available IMAGE clones when available. Each sequence has been compared to the KOG database and approximately 67% of the sequences have been assigned a putative functional category. Based on sequence homology to mouse and human, putative GO annotations have been determined. CONCLUSION: The results of the analysis have been stored in a publicly available database XenDB http://bibiserv.techfak.uni-bielefeld.de/xendb/. A unique capability of the database is the ability to batch upload cross species queries to identify potential Xenopus homologues and their associated full length clones. Examples are provided including mapping of microarray results and application of 'in silico' analysis. The ability to quickly translate the results of various species into 'Xenopus-centric' information should greatly enhance comparative embryological approaches.

Animals↗

GOToolBox: functional analysis of gene datasets based on Gene Ontology.

We have developed methods and tools based on the Gene Ontology (GO) resource allowing the identification of statistically over- or under-represented terms in a gene dataset; the clustering of functionally related genes within a set; and the retrieval of genes sharing annotations with a query gene. GO annotations can also be constrained to a slim hierarchy or a given level of the ontology. The source codes are available upon request, and distributed under the GPL license.

Animals↗

Comparison of the oxidative phosphorylation (OXPHOS) nuclear genes in the genomes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

BACKGROUND: In eukaryotic cells, oxidative phosphorylation (OXPHOS) uses the products of both nuclear and mitochondrial genes to generate cellular ATP. Interspecies comparative analysis of these genes, which appear to be under strong functional constraints, may shed light on the evolutionary mechanisms that act on a set of genes correlated by function and subcellular localization of their products. RESULTS: We have identified and annotated the Drosophila melanogaster, D. pseudoobscura and Anopheles gambiae orthologs of 78 nuclear genes encoding mitochondrial proteins involved in oxidative phosphorylation by a comparative analysis of their genomic sequences and organization. We have also identified 47 genes in these three dipteran species each of which shares significant sequence homology with one of the above-mentioned OXPHOS orthologs, and which are likely to have originated by duplication during evolution. Gene structure and intron length are essentially conserved in the three species, although gain or loss of introns is common in A. gambiae. In most tissues of D. melanogaster and A. gambiae the expression level of the duplicate gene is much lower than that of the original gene, and in D. melanogaster at least, its expression is almost always strongly testis-biased, in contrast to the soma-biased expression of the parent gene. CONCLUSIONS: Quickly achieving an expression pattern different from the parent genes may be required for new OXPHOS gene duplicates to be maintained in the genome. This may be a general evolutionary mechanism for originating phenotypic changes that could lead to species differentiation.

Animals↗

[MGAP-A microbe genome annotation platform].

A Microbe Genome Annotation Platform (MGAP) was developed and applied to the cynobacterium PCC7002 genome annotation. Various bioinformatics software tools from sequence analysis to gene identification and function prediction were implemented in MGAP. Protein sequence databases SWISSPROT and PDBseq, protein information resource InterPro and COG were also integrated in the platform. The web interface of MGAP has the functionality to display a circular map of gene distribution and GC contents throughout the genome. Detailed information such as the DNA and protein sequence, the location of genes on chromosomes can be viewed by clicking the corresponding object within the map. MGAP is based on a PC/Linux system affordable for small biological laboratories and has the advantage of using free software tools including MySQL, Apache and Perl.

Cyanobacteria↗

[High throughput screening and analysis of prostate cancer-related genes through mining databases].

BACKGROUND & OBJECTIVE: Investigation of differentially expressed genes in prostate cancer tissues may help to understand the molecular mechanism of prostate cancer and provide diagnostic markers or new targets for therapy. The Cancer Genome Anatomy Project (CGAP) public database provides an unprecedented opportunity for cancer researchers to mine genes differentially expressed in cancer tissues by bioinformatic methods. This study was to explore the feasibility of incorporating the Internet-available Serial Analysis of Gene Expression (SAGE) and cDNA databases to find human prostate cancer-related genes. METHODS: SAGE digital gene expression displayer (DGED) and cDNA DGED were used to analyze differentially expressed (>5 folds) genes in malignant prostate tissues compared with normal prostate tissues. The SAGE tags were filtered by their confidence and our 3 criteria. To test the confidence of tags of virtual digital analysis for gene identification, the modified GLGI (generation of longer cDNA fragments from SAGE tags for gene identification) was used to get the cDNA 3' end downstream of 20 tags. Main functions of all candidate genes and their relations to prostate cancer were annotated. RESULTS: Fifty-three differentially expressed genes were screened out by SAGE DGED, 26 of them were up-regulated and 27 were down-regulated in prostate cancer; 28 differentially expressed genes were got by cDNA DGED, 15 of them were up-regulated and 13 were down-regulated in prostate cancer. CONCLUSION: Reasonable use of public databases by the Internet-available tools is a simple, effective approach to get cancer-related genes, and might provide useful clues for further investigation although the results require experimental validation.

DNA, Complementary↗

Development of joint application strategies for two microbial gene finders.

MOTIVATION: As a starting point in annotation of bacterial genomes, gene finding programs are used for the prediction of functional elements in the DNA sequence. Due to the faster pace and increasing number of genome projects currently underway, it is becoming especially important to have performant methods for this task. RESULTS: This study describes the development of joint application strategies that combine the strengths of two microbial gene finders to improve the overall gene finding performance. Critica is very specific in the detection of similarity-supported genes as it uses a comparative sequence analysis-based approach. Glimmer employs a very sophisticated model of genomic sequence properties and is sensitive also in the detection of organism-specific genes. Based on a data set of 113 microbial genome sequences, we optimized a combined application approach using different parameters with relevance to the gene finding problem. This results in a significant improvement in specificity while there is similarity in sensitivity to Glimmer. The improvement is especially pronounced for GC rich genomes. The method is currently being applied for the annotation of several microbial genomes. AVAILABILITY: The methods described have been implemented within the gene prediction component of the GenDB genome annotation system.

Algorithms↗

The necessity of combining genomic and enzymatic data to infer metabolic function and pathways in the smallest bacteria: amino acid, purine and pyrimidine metabolism in Mollicutes.

Bacteria of the class Mollicutes have no cell wall. One species, Mycoplasma genitalium is the personification of the simplest form of independent cell-free life. Its small genome (580 kbp) is the smallest of any cell. Mollicutes have unique metabolic properties, perhaps because of their limited coding space and high mutability. Based on 16S rRNA analyses the Mollicutes Mycoplasma gallisepticum is thought to be the most mutable bacteria. Enzyme activities found in most Bacteria are absent from Mollicutes. The functions of apparently absent genes and enzymes can apparently be fulfilled by other genes and their expression products that have multiple capabilities. Because of these and other properties predictions of their metabolism based only on, e.g., either annotation, enzymatic assay, proteomic studies or structural analyses is problematic. To obtain a more confident appraisal of the functional capabilities of these simplest cells genomic and enzymatic data were combined to obtain a "metabolic consensus". The consensus is represented by a biochemical circuit for central metabolism involving purine and pyrimidine interconversions and their linkages to amino acid metabolism, glycolysis and the pentose phosphate pathway in three human Mollicutes pathogens: Mycoplasma pneumoniae, Mycoplasma genitalium and Ureaplasma urealyticum.

Amino Acids↗

Integrated pseudogene annotation for human chromosome 22: evidence for transcription.

Pseudogenes are inheritable genetic elements formally defined by two properties: their similarity to functioning genes and their presumed lack of activity. However, their precise characterization, particularly with respect to the latter quality, has proven elusive. An opportunity to explore this issue arises from the recent emergence of tiling-microarray data showing that intergenic regions (containing pseudogenes) are transcribed to a great degree. Here we focus on the transcriptional activity of pseudogenes on human chromosome 22. First, we integrated several sets of annotation to define a unified list of 525 pseudogenes on the chromosome. To characterize these further, we developed a comprehensive list of genomic features based on conservation in related organisms, expression evidence, and the presence of upstream regulatory sites. Of the 525 unified pseudogenes we could confidently classify 154 as processed and 49 as duplicated. Using data from tiling microarrays, especially from recent high-resolution oligonucleotide arrays, we found some evidence that up to a fifth of the 525 pseudogenes are potentially transcribed. Expressed sequence tags (EST) comparison further validated a number of these, and overall we found 17 pseudogenes with strong support for transcription. In particular, one of the pseudogenes with both EST and microarray evidence for transcription turned out to be a duplicated pseudogene in the cat eye syndrome critical region. Although we could not identify a meaningful number of transcription factor-binding sites (based on chromatin immunoprecipitation-chip data) near pseudogenes, we did find that approximately 12% of the pseudogenes had upstream CpG islands. Finally, analysis of corresponding syntenic regions in the mouse, rat and chimp genomes indicates, as previously suggested, that pseudogenes are less conserved than genes, but more preserved than the intergenic background (all notation is available from http://www.pseudogene.org).

Animals↗

Refine your search to explore more results.