Human genomic databases: a global public good?
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Development of bioinformatics tools provided researchers with the ability to identify full sets of trace element-containing proteins in organisms for which complete genomic sequences are available. Recently, independent bioinformatics methods were used to identify all, or almost all, genes encoding selenocysteine-containing proteins in human, mouse, and Drosophila genomes, characterizing entire selenoproteomes in these organisms. It also should be possible to search for entire sets of other trace element-associated proteins, such as metal-containing proteins, although methods for their identification are still in development.
Explore the source record for details and available documents.
The inherent potential underlying the sequence data produced by the International Human Genome Sequencing Consortium and other systematic sequencing projects is, obviously, tremendous. As such, it becomes increasingly important that all biologists have the ability to navigate through and cull important information from key publicly available databases. The continued rapid rise in available sequence information, particularly as model organism data is generated at breakneck speed, also underscores the necessity for all biologists to learn how to effectively make their way through the expanding "sequence information space." This review discusses some of the more commonly used tools for sequence discovery; tools have been developed for the effective and efficient mining of sequence information. These include LocusLink, which provides a gene-centric view of sequence-based information, as well as the 3 major genome browsers: the National Center for Biotechnology Information Map Viewer, the University of California Santa Cruz Genome Browser, and the European Bioinformatics Institute's Ensembl system. An overview of the types of information available through each of these front-ends is given, as well as information on tutorials and other documentation intended to increase the reader's familiarity with these tools.
We have continued to develop MITOMAP (http://www.gen.emory. edu/MITOMAP ), a comprehensive database for the human mitochondrial DNA (mtDNA). MITOMAP uses the mtDNA sequence as the unifying element for bringing together information on mitochondrial genome structure and function, pathogenic mutations and their clinical characteristics, population associated variation, and gene-gene interactions. Over the past year we have increased the degree of interlinking of MITOMAP information available on the web page, by using our generalized information management system, GENOME. As increasingly larger regions of the human genome are sequenced and characterized, the need for integrating such information is growing. Consequently, MITOMAP and GENOME provide a valuable reference for the mitochondrial biologist, in addition to being a model for the development of comprehensive, information storage and retrieval systems for other components of the human genome. This paper documents the changes to MITOMAP which have been implemented over the past year.
We have continued to develop MITOMAP, a comprehensive database for the human mitochondrial DNA (mtDNA). MITOMAP uses the mtDNA sequence as the unifying element for bringing together information on mitochondrial genome structure and function, pathogenic mutations and their clinical characteristics, population associated variation and gene-gene interactions. As increasingly larger regions of the human genome are sequenced and characterized, the need for integrating such information will grow. Consequently, MITOMAP not only provides a valuable reference for the mitochondrial biologist, it will also provide a model for the development of comprehensive, multi-media information storage and retrieval systems for other components of the human genome. This paper is an update of the changes which have occurred to MITOMAP over the past year.
The availability of a large number of sequenced microbial genomes allows us to conduct systematic studies on microbial gene regulatory systems. Computational methods, using comparative genomics approaches, are powerful tools to understand their mechanisms and evolutionary history. Recent advances in computational methodology for uncovering transcriptional regulatory components and their interactions are discussed.
BACKGROUND: Molecular maps have been developed for many species, and are of particular importance for varietal development and comparative genomics. However, despite the existence of multiple sets of linkage maps, databases of these data are lacking for many species, including peanut. DESCRIPTION: PeanutMap http://peanutgenetics.tamu.edu/cmap provides a web-based interface for viewing specific linkage groups of a map set. PeanutMap can display and compare multiple maps of a set based upon marker or trait correspondences, which is particularly important as cultivated peanut is a disomic tetraploid. The database can also compare linkage groups among multiple map sets, allowing identification of corresponding linkage groups from results of different research projects. Data from the two published peanut genome map sets, and also from three maps sets of phenotypic traits are present in the database. Data from PeanutMap have been incorporated into the Legume Information System website http://www.comparative-legumes.org to allow peanut map data to be used for cross-species comparisons. CONCLUSION: The utility of the database is expected to increase as several SSR-based maps are being developed currently, and expanded efforts for comparative mapping of legumes are underway. Optimal use of these data will benefit from the development of tools to facilitate comparative analysis.
The white-rot basidiomycete Phanerochaete chrysosporium employs extracellular enzymes to completely degrade the major polymers of wood: cellulose, hemicellulose, and lignin. Analysis of a total of 10,048 v2.1 gene models predicts 769 secreted proteins, a substantial increase over the 268 models identified in the earlier database (v1.0). Within the v2.1 'computational secretome,' 43% showed no significant similarity to known proteins, but were structurally related to other hypothetical protein sequences. In contrast, 53% showed significant similarity to known protein sequences including 87 models assigned to 33 glycoside hydrolase families and 52 sequences distributed among 13 peptidase families. When grown under standard ligninolytic conditions, peptides corresponding to 11 peptidase genes were identified in culture filtrates by mass spectrometry (LS-MS/MS). Five peptidases were members of a large family of aspartyl proteases, many of which were localized to gene clusters. Consistent with a role in dephosphorylation of lignin peroxidase, a mannose-6-phosphatase (M6Pase) was also identified in carbon-starved cultures. Beyond proteases and M6Pase, 28 specific gene products were identified including several representatives of gene families. These included 4 lignin peroxidases, 3 lipases, 2 carboxylesterases, and 8 glycosyl hydrolases. The results underscore the rich genetic diversity and complexity of P. chrysosporium's extracellular enzyme systems.
Explore the source record for details and available documents.
MOTIVATION: Sequence alignment techniques have been developed into extremely powerful tools for identifying the folding families and function of proteins in newly sequenced genomes. For a sufficiently low sequence identity it is necessary to incorporate additional structural information to positively detect homologous proteins. We have carried out an extensive analysis of the effectiveness of incorporating secondary structure information directly into the alignments for fold recognition and identification of distant protein homologs. A secondary structure similarity matrix based on a database of three-dimensionally aligned proteins was first constructed. An iterative application of dynamic programming was used which incorporates linear combinations of amino acid and secondary structure sequence similarity scores. Initially, only primary sequence information is used. Subsequently contributions from secondary structure are phased in and new homologous proteins are positively identified if their scores are consistent with the predetermined error rate. RESULTS: We used the SCOP40 database, where only PDB sequences that have 40% homology or less are included, to calibrate homology detection by the combined amino acid and secondary structure sequence alignments. Combining predicted secondary structure with sequence information results in a 8-15% increase in homology detection within SCOP40 relative to the pairwise alignments using only amino acid sequence data at an error rate of 0.01 errors per query; a 35% increase is observed when the actual secondary structure sequences are used. Incorporating predicted secondary structure information in the analysis of six small genomes yields an improvement in the homology detection of approximately 20% over SSEARCH pairwise alignments, but no improvement in the total number of homologs detected over PSI-BLAST, at an error rate of 0.01 errors per query. However, because the pairwise alignments based on combinations of amino acid and secondary structure similarity are different from those produced by PSI-BLAST and the error rates can be calibrated, it is possible to combine the results of both searches. An additional 25% relative improvement in the number of genes identified at an error rate of 0.01 is observed when the data is pooled in this way. Similarly for the SCOP40 dataset, PSI-BLAST detected 15% of all possible homologs, whereas the pooled results increased the total number of homologs detected to 19%. These results are compared with recent reports of homology detection using sequence profiling methods. AVAILABILITY: Secondary structure alignment homepage at http://lutece.rutgers.edu/ssas CONTACT: anders@rutchem.rutgers.edu; ronlevy@lutece.rutgers.edu SUPPLEMENTARY INFORMATION: Genome sequence/structure alignment results at http://lutece.rutgers.edu/ss_fold_predictions.
Zea mays DataBase (ZmDB) is a repository and analysis tool for sequence, expression and phenotype data of the major crop plant maize. The data accessible in ZmDB are mostly generated in a large collaborative project of maize gene discovery, sequencing and phenotypic analysis using a transposon tagging strategy and expressed sequence tag (EST) sequencing. ESTs constitute most of the current content. Database search tools, convenient links to external databases, and novel sequence analysis programs for spliced alignment are provided and together serve as an efficient protocol for gene discovery by sequence inspection. ZmDB can be accessed at http://zmdb. iastate.edu. ZmDB also provides web-based ordering of materials generated in the project, including EST and genomic DNA clones, seeds of mutant plants and microarrays of amplified EST and genomic DNA sequences.
The plant actin cytoskeleton is a highly dynamic, fibrous structure essential in many cellular processes including cell division and cytoplasmic streaming. This structure is stimulus responsive, being affected by internal stimuli, by biotic and abiotic stresses mediated in signal transduction pathways by actin-binding proteins. The completion of the Arabidopsis genome sequence has allowed a comparative identification of many actin-binding proteins. However, not all are conserved in plants, which possibly reflects the differences in the processes involved in morphogenesis between plant and other cells. Here we have searched for the Arabidopsis equivalents of 67 animal/fungal actin-binding proteins and show that 36 are not conserved in plants. One protein that is conserved across phylogeny is actin-depolymerizing factor or cofilin and we describe our work on the activity of vegetative tissue and pollen-specific isoforms of this protein in plant cells, concluding that they are functionally distinct.
Dictyostelium is an attractive model system for the study of mechanisms basic to cellular function or complex multicellular developmental processes. Recent advances in Dictyostelium genomics have generated a wide spectrum of resources. However, much of the current genomic sequence information is still not currently available through GenBank or related databases. Thus, many investigators are unaware that extensive sequence data from Dictyostelium has been compiled, or of its availability and access. Here, we discuss progress in Dictyostelium genomics and gene annotation, and highlight the primary portals for sequence access, manipulation and analysis (http://genome.imb-jena.de/dictyostelium/; http://dictygenome.bcm.tmc.edu/; http://www.sanger. ac.uk/Projects/D_discoideum/; http://www.csm.biol. tsukuba.ac.jp/cDNAproject.html).
MOTIVATION: In the past decade, a vast amount of mapping data has been generated on the human X chromosome, without a mechanism which would provide a global view of exactly what has been achieved. Large datasets are available electronically, but in heterogeneous formats and with incompatible access modes. In addition, relationships between objects in different datasets are often not specified. RESULTS: We discuss the problem of integrating these data into one database and define a number of requirements that are vital for any integration approach. We have developed IXDB, the Integrated X chromosome database, which fulfils those requirements and aims at providing a global view on genomic data at a chromosomal level. IXDB represents a conceptual framework based on identifying, storing and analysing relationships between biological objects, and includes a series of tools to automate the integration of such information. It currently focuses on physical mapping data, as a starting point towards a map of the human X chromosome that should provide a uniform and global research resource for ongoing and future sequencing and functional studies. AVAILABILITY: IXDB is available at http://ixdb.mpimg-berlin-dahlem.mpg.de. The iace2ixdb software and a description of the Iace data format are available from the authors. CONTACT: hrc@genoscope.cns.fr
As knowledge of human genetic polymorphisms grows, so does the opportunity and challenge of identifying those polymorphisms that may impact the health or disease risk of an individual person. A critical need is to organize large-scale polymorphism analyses and to prioritize candidate non-synonymous coding SNPs (nsSNPs) that should be tested in experimental and epidemiological studies to establish their context-specific impacts on protein function. In addition, with emerging high-resolution clinical genetics testing, new polymorphisms must be analyzed in the context of all available protein feature knowledge including other known mutations and polymorphisms. To approach this, we developed PolyDoms (http://polydoms.cchmc.org/) as a database to integrate the results of multiple algorithmic procedures and functional criteria applied to the entire Entrez dbSNP dataset. In addition to predicting structural and functional impacts of all nsSNPs, filtering functions enable group-based identification of potentially harmful nsSNPs among multiple genes associated with specific diseases, anatomies, mammalian phenotypes, gene ontologies, pathways or protein domains. PolyDoms, thus, provides a means to derive a list of candidate SNPs to be evaluated in experimental or epidemiological studies for impact on protein functions and disease risk associations. PolyDoms will continue to be curated to improve its usefulness.
PlantGDB (http://www.plantgdb.org/) is a database of molecular sequence data for all plant species with significant sequencing efforts. The database organizes EST sequences into contigs that represent tentative unique genes. Contigs are annotated and, whenever possible, linked to their respective genomic DNA. Genome sequence fragments are assembled similarly. The goal of the PlantGDB web site is to establish the basis for identifying sets of genes common to all plants or specific to particular species by integrating a number of bioinformatics tools that facilitate gene prediction and cross- species comparisons. For species with large-scale genome sequencing efforts, PlantGDB provides genome browsing capabilities that integrate all available EST and cDNA evidence for current gene models (for Arabidopsis thaliana, see the AtGDB site at http://www.plantgdb.org/AtGDB/).
The unc-47 locus of Caenorhabditis elegans has been suggested to encode a synaptic vesicle GABA transporter. Here we used hydropathy plot analysis to identify a candidate vesicular GABA transporter in genomic sequences derived from a region of the physical map comprising unc-47. A mouse homologue was identified and cloned from EST database information. In situ hybridization in rat brain revealed codistribution with both GABAergic and glycinergic neuronal markers. Moreover, expression in COS-7 and PC12 cells induced an intracellular, glycine-sensitive GABA uptake activity. These observations, consistent with previous data on GABA and glycine uptake by synaptic vesicles, demonstrate that the mouse clone encodes a vesicular inhibitory amino acid transporter.