Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

The Microbial Rosetta Stone Database: a compilation of global and emerging infectious microorganisms and bioterrorist threat agents.

BACKGROUND: Thousands of different microorganisms affect the health, safety, and economic stability of populations. Many different medical and governmental organizations have created lists of the pathogenic microorganisms relevant to their missions; however, the nomenclature for biological agents on these lists and pathogens described in the literature is inexact. This ambiguity can be a significant block to effective communication among the diverse communities that must deal with epidemics or bioterrorist attacks. RESULTS: We have developed a database known as the Microbial Rosetta Stone. The database relates microorganism names, taxonomic classifications, diseases, specific detection and treatment protocols, and relevant literature. The database structure facilitates linkage to public genomic databases. This paper focuses on the information in the database for pathogens that impact global public health, emerging infectious organisms, and bioterrorist threat agents. CONCLUSION: The Microbial Rosetta Stone is available at http://www.microbialrosettastone.com/. The database provides public access to up-to-date taxonomic classifications of organisms that cause human diseases, improves the consistency of nomenclature in disease reporting, and provides useful links between different public genomic and public health databases.

Animals↗

TcruziDB: an integrated Trypanosoma cruzi genome resource.

TcruziDB (http://TcruziDB.org) is an integrated genome database for the parasitic organism Trypanosoma cruzi, the causative agent of Chagas' disease. The database currently incorporates all available sequence data (Genomic, BAC, EST) in a single user-friendly location. The database contains a variety of tools specifically designed for searching unannotated draft sequence via BLAST, keyword searches of pre-computed BLAST results, and protein motif searches. Release 1.0 of the database contains nearly 730 million bp of genome sequence from 1.1 million sequence reads generated by the TIGR-Karolinska-SBRI Trypanosoma cruzi Genome Consortium and 15 million bp of clustered EST and genomic sequence obtained from other sources. As annotation, microarray and proteomic data become available, the database will incorporate and integrate these data using the GUS (http://www.gusdb. org) relational framework.

Animals↗

SwissRegulon: a database of genome-wide annotations of regulatory sites.

SwissRegulon (http://www.swissregulon.unibas.ch) is a database containing genome-wide annotations of regulatory sites in the intergenic regions of genomes. The regulatory site annotations are produced using a number of recently developed algorithms that operate on multiple alignments of orthologous intergenic regions from related genomes in combination with, whenever available, known sites from the literature, and ChIP-on-chip binding data. Currently SwissRegulon contains annotations for yeast and 17 prokaryotic genomes. The database provides information about the sequence, location, orientation, posterior probability and, whenever available, binding factor of each annotated site. To enable easy viewing of the regulatory site annotations in the context of other features annotated on the genomes, the sites are displayed using the GBrowse genome browser interface and can be queried based on any annotated genomic feature. The database can also be queried for regulons, i.e. sites bound by a common factor.

Algorithms↗

Molecular characterization of a 2.7-kb, 12q13-specific, retroviral-related sequence isolated by RDA from monozygotic twin pairs discordant for schizophrenia.

This report deals with the molecular characterization of a representational difference analysis (RDA)-derived sequence (SZRV-2, GenBank accession No. AF135486; Genome Database accession Nos. 7692183 and 7501402) from three monozygotic twin pairs discordant for schizophrenia (MZD). The results suggest that it is a primate-specific, heavily methylated, and placentally expressed (-7-kb mRNA) endogenous retroviral-related (ERV) sequence of the human genome. We have mapped this sequence to 12q13 using two SZRV-2 positive BAC clones (4K11 (Genome Survey Sequence Database No. 1752076; GenBank accession No. AZ301773) and 501H16) by fluorescence in situ hybridization. End sequencing of the 4K11 BAC clone has allowed identification of nearby genes from the human genome database at NCBI that may be of interest in schizophrenia research. These include viral-related sequences (potential hot spots for insertions), developmental, channel, and signal transduction genes, as well as genes affecting expression of certain receptors in neurons. Furthermore, when used as a probe on Southern blots, SZRV-2 detected no difference between schizophrenia patients from southwestern Ontario and their matched controls. However, it identified aberrant methylation in one of the eight patients and none of the 21 unaffected controls. Although additional experiments will be required to establish the significance, if any, of SZRV-2 methylation in the complex etiology of schizophrenia, molecular results included offer a novel insight into the role of retroviral-related sequences in the origin, organization, and regulation of the human genome.

Base Sequence↗

Saccharomyces cerevisiae S288C genome annotation: a working hypothesis.

The S. cerevisiae genome is the most well-characterized eukaryotic genome and one of the simplest in terms of identifying open reading frames (ORFs), yet its primary annotation has been updated continually in the decade since its initial release in 1996 (Goffeau et al., 1996). The Saccharomyces Genome Database (SGD; www.yeastgenome.org) (Hirschman et al., 2006), the community-designated repository for this reference genome, strives to ensure that the S. cerevisiae annotation is as accurate and useful as possible. At SGD, the S. cerevisiae genome sequence and annotation are treated as a working hypothesis, which must be repeatedly tested and refined. In this paper, in celebration of the tenth anniversary of the completion of the S. cerevisiae genome sequence, we discuss the ways in which the S. cerevisiae sequence and annotation have changed, consider the multiple sources of experimental and comparative data on which these changes are based, and describe our methods for evaluating, incorporating and documenting these new data.

Base Sequence↗

GeneHuggers: database mining and application connectivity tools for subsequence analyses of the human genome.

UNLABELLED: GeneHuggers is a collection of program modules that enables precise selection of subsequence regions from records of the RefSeq human genome database. Subsequence regions can be selected based on diverse criteria, including feature addresses, annotations from LocusLink and UniGene, and results obtained from analyses with homologous subsequence detection programs. GeneHuggers provides functionality to the UNIX operating system that allows customized bioinformatics program development. AVAILABILITY: GeneHuggers source code is available under the GNU general public license and can be downloaded from ftp://ftp.scripps.edu/pub/genehuggers/gh.tar.gz

Abstracting and Indexing↗

Data mining for regulatory elements in yeast genome.

We have examined methods and developed a general software tool for finding and analyzing combinations of transcription factor binding sites that occur relatively often in gene upstream regions (putative promoter regions) in the yeast genome. Such frequently occurring combinations may be essential parts of possible promoter classes. The regions upstream to all genes were first isolated from the yeast genome database MIPS using the information in the annotation files of the database. The ones that do not overlap with coding regions were chosen for further studies. Next, all occurrences of the yeast transcription factor binding sites, as given in the IMD database, were located in the genome and in the selected regions in particular. Finally, by using a general purpose data mining software in combination with our own software, which parametrizes the search, we can find the combinations of binding sites that occur in the upstream regions more frequently than would be expected on the basis of the frequency of individual sites. The procedure also finds so-called association rules present in such combinations. The developed tool is available for use through the WWW.

Binding Sites↗

Real-World Treatment Patterns and Clinical Outcomes After First-Line Therapy in Patients with KRAS G12C-Mutant Advanced Non-Small-Cell Lung Cancer in the United States.

BACKGROUND: Approximately 13% of NSCLC cases have KRAS G12C mutations. As therapeutic strategies targeting KRAS G12C-mutant NSCLC evolve, it is important to understand clinical presentation and current outcomes for these patients. METHODS: This retrospective study used data from two US nationwide databases, an electronic health records (EHR) database and a clinico-genomic database (CGDB) of EHR data linked to data from comprehensive genomic profiling tests. Eligible patients had advanced NSCLC, initiated first-line therapy from August 2018 to December 2022, and had KRAS test results. Clinicopathologic characteristics, treatments, real-world progression-free survival (rwPFS), and overall survival (OS) were analyzed. RESULTS: There were 1227 patients with KRAS G12C-mutant NSCLC in the EHR database and 447 in the CGDB. First-line regimen was platinum-based chemotherapy plus pembrolizumab for 46% and pembrolizumab monotherapy for 20%. Less than 40% of patients received second-line therapy. Median (95% CI) OS for KRAS G12C-mutant NSCLC patients in the EHR was 17.0 (15.2-18.9) months. Variables significantly associated with shorter OS included PD-L1 <1%, brain metastases, STK11 co-mutation, and poor performance status. Patients treated with platinum-based chemotherapy plus pembrolizumab had median rwPFS of 5.3 (4.5-7.3) months and OS of 12.8 (11.1-17.3) months in the CGDB; median OS was 15.6 (12.5-18.6) months in the EHR. Patients with PD-L1 &#x2265; 50% treated with pembrolizumab monotherapy had median rwPFS of 4.6 (3.0-15.6) months and OS of 20.4 (10.3-38.5) months in the CGDB; median OS was 22.1 (18.7-30.7) in the EHR. CONCLUSIONS: These data provide a real-world benchmark of outcomes for patients with KRAS G12C-mutant NSCLC receiving the current standard of care and indicate an unmet need for more effective first-line therapies.

KRAS G12C↗

An interactive bovine in silico SNP database (IBISS).

An interactive bovine in silico SNP (IBISS) database has been created through the clustering and aligning of bovine EST and mRNA sequences. Approximately 324,000 EST and mRNA sequences were clustered to produce 29,965 clusters (producing 48,679 consensus sequences) and 48,565 singletons. A SNP screening regime was placed on variations detected in the multiple sequence alignment files to determine which SNPs are more likely to be real rather than sequencing errors. A small subset of predicted SNPs was validated on a diverse set of bovine DNA samples using PCR amplification and sequencing. Fifty percent of the predicted SNPs in the "putative >1" category were polymorphic in the population sampled. The IBISS database represents more than just a SNP database; it is also a genomic database containing uniformly annotated predicted gene mRNA and protein sequences, gene structure, and genomic organization information.

Animals↗

Brassica ASTRA: an integrated database for Brassica genomic research.

Brassica ASTRA is a public database for genomic information on Brassica species. The database incorporates expressed sequences with Swiss-Prot and GenBank comparative sequence annotation as well as secondary Gene Ontology (GO) annotation derived from the comparison with Arabidopsis TAIR GO annotations. Simple sequence repeat molecular markers are identified within resident sequences and mapped onto the closely related Arabidopsis genome sequence. Bacterial artificial chromosome (BAC) end sequences derived from the Multinational Brassica Genome Project are also mapped onto the Arabidopsis genome sequence enabling users to identify candidate Brassica BACs corresponding to syntenic regions of Arabidopsis. This information is maintained in a MySQL database with a web interface providing the primary means of interrogation. The database is accessible at http://hornbill.cspp.latrobe.edu.au.

Brassica↗

IMGT-ONTOLOGY and IMGT databases, tools and Web resources for immunogenetics and immunoinformatics.

The international ImMunoGeneTics information system (IMGT; http://imgt.cines.fr), is a high quality integrated information system specialized in immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC), and related proteins of the immune system (RPI) of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM; Université Montpellier II and CNRS) at Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT consists of several sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB) and one 3D structure database (IMGT/3Dstructure-DB), interactive tools for sequence analysis (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/PhyloGene, IMGT/Allele-Align), for genome analysis (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView) and for 3D structure analysis (IMGT/StructuralQuery), and Web resources ("IMGT Marie-Paule page") comprising 8000 HTML pages. IMGT other accesses include SRS, FTP, search by BLAST, etc. By its high quality and its easy data distribution, IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), veterinary research, genome diversity and genome evolution studies of the adaptive immune responses, biotechnology related to antibody engineering (single chain Fragment variable (scFv), phage displays, combinatorial libraries) and therapeutical approaches (grafts, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

IMGT-ONTOLOGY for immunogenetics and immunoinformatics.

IMGT, the international ImMunoGeneTics information system(R) (http://imgt.cines.fr), is a high quality integrated knowledge resource specializing in immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC) and related proteins of the immune system (RPI) of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGenetique Moleculaire LIGM. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT consists of several sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB) and one three-dimensional structure database (IMGT/3Dstructure-DB), interactive tools for sequence analysis (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/PhyloGene, IMGT/Allele-Align), for genome analysis (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView) and for 3D structure analysis (IMGT/StructuralQuery), and Web resources ("IMGT Marie-Paule page") comprising 8000 HTML pages. IMGT other accesses include SRS, FTP, search by BLAST, etc. By its high quality and its easy data distribution, IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), veterinary research, genome diversity and genome evolution studies of the adaptive immune responses, biotechnology related to antibody engineering (scFv, phage displays, combinatorial libraries) and therapeutical approaches (grafts, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

ZmDB, an integrated database for maize genome research.

Zea mays DataBase (ZmDB) seeks to provide a comprehensive view of maize (corn) genetics by linking genomic sequence data with gene expression analysis and phenotypes of mutant plants. ZmDB originated in 1999 as the Web portal for a large project of maize gene discovery, sequencing and phenotypic analysis using a transposon tagging strategy and expressed sequence tag (EST) sequencing. Recently, ZmDB has broadened its scope to include all public maize ESTs, genome survey sequences (GSSs), and protein sequences. More than 170 000 ESTs are currently clustered into approximately 20 000 contigs and about an equal number of apparent singlets. These clusters are continuously updated and annotated with respect to potential encoded protein products. More than 100 000 GSSs are similarly assembled and annotated by spliced alignment with EST and protein sequences. The ZmDB interface provides quick access to analytical tools for further sequence analysis. Every sequence record is linked to several display options and similarity search tools, including services for multiple sequence alignment, protein domain determination and spliced alignment. Furthermore, ZmDB provides web-based ordering of materials generated in the project, including ESTs, ordered collections of genomic sequences tagged with the RescueMu transposon and microarrays of amplified ESTs. ZmDB can be accessed at http://zmdb.iastate.edu/.

DNA Transposable Elements↗

IMGT, the international ImMunoGeneTics database.

The international ImMunoGeneTics database (IMGT) (http://imgt.cines.fr), is a high quality integrated information system specializing in Immunoglobulins (IG), T cell Receptors (TR) and Major Histocompatibility Complex (MHC) of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB) with different interfaces (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ('IMGT Marie-Paule page') and interactive tools for sequence analysis (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene). IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on IMGT-ONTOLOGY. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and thera-peutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

Perspective: the ovarian kaleidoscope database-II. Functional genomic analysis of an organ-specific database.

In the postgenomic era, it is now possible to investigate the function of all human genes to provide an integrated view of physiology and pathophysiology. An organ-based approach has been used to set up a database integrating existing text-based literature on individual ovarian genes and their sequence-based data in the GenBank. The Ovarian Kaleidoscope database (OKdb) has accumulated nearly one thousand individual gene pages that are searchable based on gene function, cellular localization, chromosomal position, ovarian cell type, ovarian function, mutant phenotypes, and other criteria. The present review exemplifies the use of this organ-based database in setting up gene pathway maps for DNA array analysis, identifying key gene networks essential for infertility phenotypes, comparing chromosomal synteny regions for finding candidate fertility genes, categorizing cell-specific and hormonally coregulated genes for promoter analysis, and documenting potential ligands and receptors in the paracrine regulation of follicular development. The present global analysis of gene function and relationships in an organ-specific manner provides a functional genomic paradigm for the future understanding of the physiology and pathophysiology of diverse organs.

Animals↗

CpG islands of the pig.

We describe an analysis of the CpG islands (CGIs) of the pig. We have used both database survey and a porcine genomic library that is enriched for CGIs. Approximately half of 41 pig genomic database sequences had CGIs with an average G + C content of 65.3%, an average CpG observed/expected frequency of 0.85, and an average size of 978 bp. Of 27 CGI library clones, 16 were nonrepetitive, nonribosomal DNA and CGI-like. CGI library clones had similar average values for G + C and CpG frequency to CGIs of database genes, and an average size of 670 bp, as MseI cuts within some islands. Library clones were also shown to be low copy number and unmethylated in genomic DNA. The presence in the library of seven previously known CGI sequences was confirmed as was the absence of one nonisland sequence. The CGI library exhibits an R-band pattern for many chromosomes in FISH analysis. The pig chromosome arms that show the most dense CGI population are homologous to segments of human chromosomes that are known to be gene rich.

Animals↗

Functional and structural genomics using PEDANT.

MOTIVATION: Enormous demand for fast and accurate analysis of biological sequences is fuelled by the pace of genome analysis efforts. There is also an acute need in reliable up-to-date genomic databases integrating both functional and structural information. Here we describe the current status of the PEDANT software system for high-throughput analysis of large biological sequence sets and the genome analysis server associated with it. RESULTS: The principal features of PEDANT are: (i) completely automatic processing of data using a wide range of bioinformatics methods, (ii) manual refinement of annotation, (iii) automatic and manual assignment of gene products to a number of functional and structural categories, (iv) extensive hyperlinked protein reports, and (v) advanced DNA and protein viewers. The system is easily extensible and allows to include custom methods, databases, and categories with minimal or no programming effort. PEDANT is actively used as a collaborative environment to support several on-going genome sequencing projects. The main purpose of the PEDANT genome database is to quickly disseminate well-organized information on completely sequenced and unfinished genomes. It currently includes 80 genomic sequences and in many cases serves as the only source of exhaustive information on a given genome. The database also acts as a vehicle for a number of research projects in bioinformatics. Using SQL queries, it is possible to correlate a large variety of pre-computed properties of gene products encoded in complete genomes with each other and compare them with data sets of special scientific interest. In particular, the availability of structural predictions for over 300 000 genomic proteins makes PEDANT the most extensive structural genomics resource available on the web.

Arabidopsis↗

GNARE: automated system for high-throughput genome analysis with grid computational backend.

Recent progress in genomics and experimental biology has brought exponential growth of the biological information available for computational analysis in public genomics databases. However, applying the potentially enormous scientific value of this information to the understanding of biological systems requires computing and data storage technology of an unprecedented scale. The Grid, with its aggregated and distributed computational and storage infrastructure, offers an ideal platform for high-throughput bioinformatics analysis. To leverage this we have developed the Genome Analysis Research Environment (GNARE)--a scalable computational system for the high-throughput analysis of genomes, which provides an integrated database and computational backend for data-driven bioinformatics applications. GNARE efficiently automates the major steps of genome analysis including acquisition of data from multiple genomic databases; data analysis by a diverse set of bioinformatics tools; and storage of results and annotations. High-throughput computations in GNARE are performed using distributed heterogeneous Grid computing resources such as Grid2003, TeraGrid, and the DOE Science Grid. Multi-step genome analysis workflows involving massive data processing, the use of application-specific tools and algorithms and updating of an integrated database to provide interactive web access to results are all expressed and controlled by a "virtual data" model which transparently maps computational workflows to distributed Grid resources. This paper describes how Grid technologies such as Globus, Condor, and the Gryphyn Virtual Data System were applied in the development of GNARE. It focuses on our approach to Grid resource allocation and to the use of GNARE as a computational framework for the development of bioinformatics applications.

Computational Biology↗