Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

The phytophthora genome initiative database: informatics and analysis for distributed pathogenomic research.

The Phytophthora Genome Initiative (PGI) is a distributed collaboration to study the genome and evolution of a particularly destructive group of plant pathogenic oomycete, with the goal of understanding the mechanisms of infection and resistance. NCGR provides informatics support for the collaboration as well as a centralized data repository. In the pilot phase of the project, several investigators prepared Phytophthora infestans and Phytophthora sojae EST and Phytophthora sojae BAC libraries and sent them to another laboratory for sequencing. Data from sequencing reactions were transferred to NCGR for analysis and curation. An analysis pipeline transforms raw data by performing simple analyses (i.e., vector removal and similarity searching) that are stored and can be retrieved by investigators using a web browser. Here we describe the database and access tools, provide an overview of the data therein and outline future plans. This resource has provided a unique opportunity for the distributed, collaborative study of a genus from which relatively little sequence data are available. Results may lead to insight into how better to control these pathogens. The homepage of PGI can be accessed at http:www.ncgr.org/pgi, with database access through the database access hyperlink.

Databases, Factual↗

Cloning and genomic localization of the murine LPS-induced CXC chemokine (LIX) gene, Scyb5.

LPS-induced CXC chemokine (LIX) is a murine chemokine similar to two human chemokines, ENA-78 (CXCL5) and GCP-2 (CXCL6). To clarify the relationship of LIX to human ENA-78 and GCP-2, we cloned and mapped the LIX gene. The organization of the LIX gene ( Scyb5) is similar to those of the human ENA-78 ( SCYB5) and GCP-2 ( SCYB6) genes. The intron-exon boundaries of the three genes are exactly conserved, and the introns have similar sizes. The first 100 bp of the 5' flanking regions are highly similar, with conserved NF-kappaB and GATA sites in identical positions in all three genes. Further 5', the Lix flanking region sequence diverges from those of ENA-78 and GCP-2, which remain highly similar for 350 bp preceding the start sites. Using a (C57BL/6 J x Mus spretus) F1 x C57BL/6J backcross panel, Lix was mapped to a locus near D5Ucla5 at 49.0 cM on Chromosome (Chr) 5. Mapping with the T31 radiation hybrid panel placed Lix between D5Mit360 and D5Mit6. Physical maps of the CXC chemokine clusters on murine Chr 5 and human Chr 21 were constructed using the Celera mouse genome database and the public human genome database. The sequence and mapping data suggest that the human ENA78-PBP-PF4 and GCP2- psi PBP-PF4V1 loci arose from an evolutionarily recent duplication of an ancestral locus related to the murine Lix-Pbp-Pf4 locus.

Animals↗

The PRESAGE database for structural genomics.

The PRESAGE database is a collaborative resource for structural genomics. It provides a database of proteins to which researchers add annotations indicating current experimental status, structural predictions and suggestions. The database is intended to enhance communication among structural genomics researchers and aid dissemination of their results. The PRESAGE database may be accessed at http://presage.stanford.edu/

Databases, Factual↗

An integrated database of the ascidian, Ciona intestinalis: towards functional genomics.

An integrated genome database is essential for future studies of functional genomics. In this study, we update cDNA and genomic resources of the ascidian, Ciona intestinalis, and provide an integrated database of the genomic and cDNA data by extending a database published previously. The updated resources include over 190,000 ESTs (672,396 in total together with the previous ESTs) and over 1,000 full-insert sequences (6,773 in total). In addition, results of mapping information of the determined scaffolds onto chromosomes, ESTs from a full-length enriched cDNA library for indication of precise 5'-ends of genes, and comparisons of SNPs and indels among different individuals are integrated into this database, all of these results being reported recently. These advances continue to increase the utility of Ciona intestinalis as a model organism whilst the integrated database will be useful for researchers in comparative and evolutionary genomics.

Animals↗

HuGeMap: a distributed and integrated Human Genome Map database.

The HuGeMap database stores the major genetic and physical maps of the human genome. It is also interconnected with the gene radiation hybrid mapping database RHdb. HuGeMap is accessible through a Web server for interactive browsing at URL http://www.infobiogen. fr/services/Hugemap , as well as through a CORBA server for effective programming. HuGeMap is intended as an attempt to build open, interconnected databases, that is databases that distribute their objects worldwide in compliance with a recognized standard of distribution. Maps can be displayed and compared with a java applet (http://babbage.infobiogen.fr:15000/Mappet/Show. html ) that queries the HuGeMap ORB server as well as the RHdb ORB server at the EBI.

Chromosome Mapping↗

CBS Genome Atlas Database: a dynamic storage for bioinformatic results and sequence data.

UNLABELLED: Currently, new bacterial genomes are being published on a monthly basis. With the growing amount of genome sequence data, there is a demand for a flexible and easy-to-maintain structure for storing sequence data and results from bioinformatic analysis. More than 150 sequenced bacterial genomes are now available, and comparisons of properties for taxonomically similar organisms are not readily available to many biologists. In addition to the most basic information, such as AT content, chromosome length, tRNA count and rRNA count, a large number of more complex calculations are needed to perform detailed comparative genomics. DNA structural calculations like curvature and stacking energy, DNA compositions like base skews, oligo skews and repeats at the local and global level are just a few of the analysis that are presented on the CBS Genome Atlas Web page. Complex analysis, changing methods and frequent addition of new models are factors that require a dynamic database layout. Using basic tools like the GNU Make system, csh, Perl and MySQL, we have created a flexible database environment for storing and maintaining such results for a collection of complete microbial genomes. Currently, these results counts to more than 220 pieces of information. The backbone of this solution consists of a program package written in Perl, which enables administrators to synchronize and update the database content. The MySQL database has been connected to the CBS web-server via PHP4, to present a dynamic web content for users outside the center. This solution is tightly fitted to existing server infrastructure and the solutions proposed here can perhaps serve as a template for other research groups to solve database issues. AVAILABILITY: A web based user interface which is dynamically linked to the Genome Atlas Database can be accessed via www.cbs.dtu.dk/services/GenomeAtlas/. SUPPLEMENTARY INFORMATION: This paper has a supplemental information page which links to the examples presented: www.cbs.dtu.dk/services/GenomeAtlas/suppl/bioinfdatabase.

Algorithms↗

IMGT/GENE-DB: a comprehensive database for human and mouse immunoglobulin and T cell receptor genes.

IMGT/GENE-DB is the comprehensive IMGT genome database for immunoglobulin (IG) and T cell receptor (TR) genes from human and mouse, and, in development, from other vertebrates. IMGT/GENE-DB is the international reference for the IG and TR gene nomenclature and works in close collaboration with the HUGO Nomenclature Committee, Mouse Genome Database and genome committees for other species. IMGT/GENE-DB allows a search of IG and TR genes by locus, group and subgroup, which are CLASSIFICATION concepts of IMGT-ONTOLOGY. Short cuts allow the retrieval gene information by gene name or clone name. Direct links with configurable URL give access to information usable by humans or programs. An IMGT/GENE-DB entry displays accurate gene data related to genome (gene localization), allelic polymorphisms (number of alleles, IMGT reference sequences, functionality, etc.) gene expression (known cDNAs), proteins and structures (Protein displays, IMGT Colliers de Perles). It provides internal links to the IMGT sequence databases and to the IMGT Repertoire Web resources, and external links to genome and generalist sequence databases. IMGT/GENE-DB manages the IMGT reference directory used by the IMGT tools for IG and TR gene and allele comparison and assignment, and by the IMGT databases for gene data annotation. IMGT/GENE-DB is freely available at http://imgt.cines.fr.

Alleles↗

Growth hormone transcription factor ZN-16 genomic coding regions are composed of a single exon and are evolutionarily conserved in mammals.

The structure of the gene encoding ZN-16, a transcription factor that binds to the mammalian growth hormone promoter in tandem with Pit-1, was determined in order to elucidate the exon-intron organization of the 16 zinc finger domains of the protein. Southern hybridization of mouse genomic DNA showed fragments with sizes identical to those predicted from mouse ZN-16 cDNA for two different probes covering the 2200 aa coding frame. Mouse genome database sequences also showed no introns in the zn-16 coding regions on chromosome 4. Analysis of human zn-16 by Southern hybridization and genomic database sequence analysis also indicated a single exon for the human protein coding sequences. BLASTP query of available genomic databases with critical zinc finger residues from mouse ZN-16 identified highly similar canine, bovine, and chimpanzee genomic sequences that encode proteins. Phylogenetic analysis of these mammalian proteins resulted in relationships as would be expected in species spanning rodents to humans. All six independent zn-16 sequences show a single exon coding region with no introns, a similarity ruling out the possibility that these genomic sequences are pseudogenes. Thus, mammalian zn-16 has a compact single exon structure encoding a very large protein (2200-3000 aa), the conservation of which may have functional implications such as the importance of posttranscriptional modifications.

Amino Acid Sequence↗

A SNP-centric database for the investigation of the human genome.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are an increasingly important tool for genetic and biomedical research. Although current genomic databases contain information on several million SNPs and are growing at a very fast rate, the true value of a SNP in this context is a function of the quality of the annotations that characterize it. Retrieving and analyzing such data for a large number of SNPs often represents a major bottleneck in the design of large-scale association studies. DESCRIPTION: SNPper is a web-based application designed to facilitate the retrieval and use of human SNPs for high-throughput research purposes. It provides a rich local database generated by combining SNP data with the Human Genome sequence and with several other data sources, and offers the user a variety of querying, visualization and data export tools. In this paper we describe the structure and organization of the SNPper database, we review the available data export and visualization options, and we describe how the architecture of SNPper and its specialized data structures support high-volume SNP analysis. CONCLUSIONS: The rich annotation database and the powerful data manipulation and presentation facilities it offers make SNPper a very useful online resource for SNP research. Its success proves the great need for integrated and interoperable resources in the field of computational biology, and shows how such systems may play a critical role in supporting the large-scale computational analysis of our genome.

Databases, Genetic↗

SGMD: the Soybean Genomics and Microarray Database.

The Soybean Genomics and Microarray Database (SGMD) attempts to provide an integrated view of the interaction of soybean with the soybean cyst nematode and contains genomic, EST and microarray data with embedded analytical tools allowing correlation of soybean ESTs with their gene expression profiles. SGMD provides analytical tools to mine the microarray data quickly by integrating many analysis methods within the database itself. The expression profiles of genes at time intervals during the first 8 days of nematode invasion is searchable by gene name or GenBank accession number. Recent developments include the addition of a searchable database for soybean cyst nematode ESTs and photographs of the invasion process at time points examined using microarrays. SGMD is completely accessible from the web at: http://psi081.ba.ars.usda.gov/SGMD/default.htm.

Computational Biology↗

In silico reconstruction of the metabolic pathways of Lactobacillus plantarum: comparing predictions of nutrient requirements with those from growth experiments.

On the basis of the annotated genome we reconstructed the metabolic pathways of the lactic acid bacterium Lactobacillus plantarum WCFS1. After automatic reconstruction by the Pathologic tool of Pathway Tools (http://bioinformatics.ai.sri.com/ptools/), the resulting pathway-genome database, LacplantCyc, was manually curated extensively. The current database contains refinements to existing routes and new gram-positive bacterium-specific reactions that were not present in the MetaCyc database. These reactions include, for example, reactions related to cell wall biosynthesis, molybdopterin biosynthesis, and transport. At present, LacplantCyc includes 129 pathways and 704 predicted reactions involving some 670 chemical species and 710 enzymes. We tested vitamin and amino acid requirements of L. plantarum experimentally and compared the results with the pathways present in LacplantCyc. In the majority of cases (32 of 37 cases) the experimental results agreed with the final reconstruction. LacplantCyc is the most extensively curated pathway-genome database for gram-positive bacteria and is open to the microbiology community via the World Wide Web (www.lacplantcyc.nl). It can be used as a reference pathway-genome database for gram-positive microbes in general and lactic acid bacteria in particular.

Amino Acids↗

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequences representing genomic data, transcripts and proteins. Although the goal is to provide a comprehensive dataset representing the complete sequence information for any given species, the database pragmatically includes sequence data that are currently publicly available in the archival databases. The database incorporates data from over 2400 organisms and includes over one million proteins representing significant taxonomic diversity spanning prokaryotes, eukaryotes and viruses. Nucleotide and protein sequences are explicitly linked, and the sequences are linked to other resources including the NCBI Map Viewer and Gene. Sequences are annotated to include coding regions, conserved domains, variation, references, names, database cross-references, and other features using a combined approach of collaboration and other input from the scientific community, automated annotation, propagation from GenBank and curation by NCBI staff.

Animals↗

The insertional history of an active family of L1 retrotransposons in humans.

As humans contain a currently active L1 (LINE-1) non-LTR retrotransposon family (Ta-1), the human genome database likely provides only a partial picture of Ta-1-generated diversity. We used a non-biased method to clone Ta-1 retrotransposon-containing loci from representatives of four ethnic populations. We obtained 277 distinct Ta-1 loci and identified an additional 67 loci in the human genome database. This collection represents approximately 90% of the Ta-1 population in the individuals examined and is thus more representative of the insertional history of Ta-1 than the human genome database, which lacked approximately 40% of our cloned Ta-1 elements. As both polymorphic and fixed Ta-1 elements are as abundant in the GC-poor genomic regions as in ancestral L1 elements, the enrichment of L1 elements in GC-poor areas is likely due to insertional bias rather than selection. Although the chromosomal distribution of Ta-1 inserts is generally a function of chromosomal length and gene density, chromosome 4 significantly deviates from this pattern and has been much more hospitable to Ta-1 insertions than any other chromosome. Also, the intra-chromosomal distribution of Ta-1 elements is not uniform. Ta-1 elements tend to cluster, and the maximal gaps between Ta-1 inserts are larger than would be expected from a model of uniform random insertion.

Chromosome Mapping↗

Functional inferences from reconstructed evolutionary biology involving rectified databases--an evolutionarily grounded approach to functional genomics.

If bioinformatics tools are constructed to reproduce the natural, evolutionary history of the biosphere, they offer powerful approaches to some of the most difficult tasks in genomics, including the organization and retrieval of sequence data, the updating of massive genomic databases, the detection of database error, the assignment of introns, the prediction of protein conformation from protein sequences, the detection of distant homologs, the assignment of function to open reading frames, the identification of biochemical pathways from genomic data, and the construction of a comprehensive model correlating the history of biomolecules with the history of planet Earth.

Amino Acid Sequence↗

The Genomes On Line Database (GOLD) v.2: a monitor of genome projects worldwide.

The Genomes On Line Database (GOLD) is a web resource for comprehensive access to information regarding complete and ongoing genome sequencing projects worldwide. The database currently incorporates information on over 1500 sequencing projects, of which 294 have been completed and the data deposited in the public databases. GOLD v.2 has been expanded to provide information related to organism properties such as phenotype, ecotype and disease. Furthermore, project relevance and availability information is now included. GOLD is available at http://www.genomesonline.org. It is also mirrored at the Institute of Molecular Biology and Biotechnology, Crete, Greece at http://gold.imbb.forth.gr/

Databases, Nucleic Acid↗

Characterization of isoforms and genomic organization of mouse calumenin.

Calumenin is a multiple EF-hand protein located in endo/sarcoplasmic reticulum of mammalian heart and other tissues [J. Biol. Chem. 272 (1997) 18232; Genomics 49 (1998) 331; Biochim. Biophys. Acta 1386 (1998) 121]. In the present study, a new isoform of mouse calumenin (mouse calumenin 2) was cloned by RT-PCR and genomic DNA PCR. The deduced amino acid sequence of mouse calumenin 2 is 315 aa long with the calculated MW of 37,064 and pI of 4.26. It has 92% aa sequence identity to previously identified mouse calumenin [J. Biol. Chem. 272 (1997) 18232] (mouse calumenin 1). The difference in the aa sequence was restricted to the first two EF-hand regions (residues 74-138). Northern blot analysis shows that mouse calumenin 2 is highly expressed in heart, lung, testis and unpregnant uterus. The expression of mouse calumenin 2 appears to decrease when fetal development is progressed. Genomic DNA PCR, sequencing and data mining of mouse genome database were utilized to examine the exon-intron boundaries of mouse calumenin genes. Both mouse calumenin 1 and 2 genes encompass six exons, and five of them (Exon1, 3, 4, 5 and 6) are identical. However, mouse calumenin 1 contains Exon2-1, whereas mouse calumenin 2 contains a neighboring Exon2-2. The calumenin genes are localized on mouse chromosome 6 having conserved synteny with human chromosome 7q32. For comparison, the genomic organization of human calumenin was also examined using the published human genome database (UCSC Genome Bioinformatics at ). Like mouse calumenin genes, two human calumenin genes also consist of five identical exons (Exon1, 3, 4, 5 and 6) and a different Exon2. The present study suggests that the genomic organization of calumenin genes is well conserved between human and mouse.

Amino Acid Sequence↗

MIPSPlantsDB--plant database resource for integrative and comparative plant genome research.

Genome-oriented plant research delivers rapidly increasing amount of plant genome data. Comprehensive and structured information resources are required to structure and communicate genome and associated analytical data for model organisms as well as for crops. The increase in available plant genomic data enables powerful comparative analysis and integrative approaches. PlantsDB aims to provide data and information resources for individual plant species and in addition to build a platform for integrative and comparative plant genome research. PlantsDB is constituted from genome databases for Arabidopsis, Medicago, Lotus, rice, maize and tomato. Complementary data resources for cis elements, repetive elements and extensive cross-species comparisons are implemented. The PlantsDB portal can be reached at http://mips.gsf.de/projects/plants.

Arabidopsis↗