Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

EcoGene: a genome sequence database for Escherichia coli K-12.

The EcoGene database provides a set of gene and protein sequences derived from the genome sequence of Escherichia coli K-12. EcoGene is a source of re-annotated sequences for the SWISS-PROT and Colibri databases. EcoGene is used for genetic and physical map compilations in collaboration with the Coli Genetic Stock Center. The EcoGene12 release includes 4293 genes. EcoGene12 differs from the GenBank annotation of the complete genome sequence in several ways, including (i) the revision of 706 predicted or confirmed gene start sites, (ii) the correction or hypothetical reconstruction of 61 frame-shifts caused by either sequence error or mutation, (iii) the reconstruction of 14 protein sequences interrupted by the insertion of IS elements, and (iv) pre-dictions that 92 genes are partially deleted gene fragments. A literature survey identified 717 proteins whose N-terminal amino acids have been verified by sequencing. 12 446 cross-references to 6835 literature citations and s are provided. EcoGene is accessible at a new website: http://bmb.med.miami.edu/EcoGene/EcoWeb. Users can search and retrieve individual EcoGene GenePages or they can download large datasets for incorporation into database management systems, facilitating various genome-scale computational and functional analyses.

Databases, Factual↗

[Introduction to Go! Poly, a human genome polymorphism database].

Databases play an important role in the study of genetic polymorphism. To meet the need for more studies of human genome polymorphism by Chinese medical and pharmaceutical community, a gene oriented human genome polymorphism database-Go! Poly was constructed. As a generalized polymorphism database, Go! Poly extracted human gene-linked sequence variations of all common types from various public resources including scientific journals and Web resources such as HGBASE (http://hgbase.cgr.ki.se) and dbSNP (http://www.ncbi.nlm.nih.gov/SNP/). The polymorphism data were then categorized into different gene loci, and the reference sequences given by LocusLink were used as positioning reference. To facilitate the use, a friendly web interface and a text based query strategy were implemented. Users can fetch specific polymorphism data in just three steps: find specific gene locus by simple search, display sequence variation information of a specific gene locus select, and view the final result of a specific variation site. Besides, a web-based submission tool is provided for direct submission, which can make the polymorphism information generated by the Chinese scientific community available from this resource.

China↗

Individual metabolism should guide agriculture toward foods for improved health and nutrition.

Genomics and bioinformatics have the vast potential to identify genes that cause disease by investigating whole-genome databases. Comparison of an individual's geno-type with a genomic database will allow the prescription of drugs to be tailored to an individual's genotype. This same bioinformatic approach, applied to the study of human metabolites, has the potential to identify and validate targets to improve personalized nutritional health and thus serve to define the added value for the next generation of foods and crops. Advances in high-throughput analytic chemistry and computing technologies make the creation of a vast database of metabolites possible for several subsets of metabolites, including lipids and organic acids. In creating integrative databases of metabolites for bioinformatic investigation, the current concept of measuring single biomarkers must be expanded to 3 dimensions to 1) include a highly comprehensive set of metabolite measurements (a profile) by multiparallel analyses, 2) measure the metabolic profile of individuals over time rather than simply in the fasted state, and 3) integrate these metabolic profiles with genomic, expression, and proteomic databases. Application of the knowledge of individual metabolism will revolutionize the ability of nutrition to deliver health benefits through food in the same way that knowledge of genomics will revolutionize individual treatment of dis-ease with pharmaceuticals.

Biomarkers↗

Genomic pathways database and biological data management.

In this paper, we discuss the properties of biological data and challenges it poses for data management, and argue that, in order to meet the data management requirements for 'digital biology', careful integration of the existing technologies and the development of new data management techniques for biological data are needed. Based on this premise, we present PathCase: Case Pathways Database System. PathCase is an integrated set of software tools for modelling, storing, analysing, visualizing and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (i) genomic information integrated with other biological data and presented starting from pathways; (ii) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (iii) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels and to pose exploratory queries; (iv) a wide range of different types of queries including, 'path' and 'neighbourhood queries' and graphical visualization of query outputs; and (v) an implementation that allows for web (XML)-based dissemination of query outputs (i.e. pathways data in BIOPAX format) to researchers in the community, giving them control on the use of pathways data.

Computational Biology↗

Informatics for mouse genetics and genome mapping.

Bioinformatics has become an essential part of biological research. The rapid pace of technology development and the ability to carry out biological experimentation in large scale require computerized systems for data management, analysis, and display. Experimentation with the mouse, a major model organism of the Human Genome Initiative, has intensified the need for bioinformatics tools for mouse mapping and genome analysis. This article describes the Mouse Genome Database in the United States, a primary resource for mouse genomic data, as well as resources at the Mammalian Genetics Unit in the United Kingdom and the Animal Genome Database of Japan. Internet addresses are provided for major genetic and physical mapping resources, major genome data sites, and resources of molecular information.

Animals↗

Expansion of the BioCyc collection of pathway/genome databases to 160 genomes.

The BioCyc database collection is a set of 160 pathway/genome databases (PGDBs) for most eukaryotic and prokaryotic species whose genomes have been completely sequenced to date. Each PGDB in the BioCyc collection describes the genome and predicted metabolic network of a single organism, inferred from the MetaCyc database, which is a reference source on metabolic pathways from multiple organisms. In addition, each bacterial PGDB includes predicted operons for the corresponding species. The BioCyc collection provides a unique resource for computational systems biology, namely global and comparative analyses of genomes and metabolic networks, and a supplement to the BioCyc resource of curated PGDBs. The Omics viewer available through the BioCyc website allows scientists to visualize combinations of gene expression, proteomics and metabolomics data on the metabolic maps of these organisms. This paper discusses the computational methodology by which the BioCyc collection has been expanded, and presents an aggregate analysis of the collection that includes the range of number of pathways present in these organisms, and the most frequently observed pathways. We seek scientists to adopt and curate individual PGDBs within the BioCyc collection. Only by harnessing the expertise of many scientists we can hope to produce biological databases, which accurately reflect the depth and breadth of knowledge that the biomedical research community is producing.

Animals↗

Oryzabase. An integrated biological and genome information database for rice.

The aim of Oryzabase is to create a comprehensive view of rice (Oryza sativa) as a model monocot plant by integrating biological data with molecular genomic information (http://www.shigen.nig.ac.jp/rice/oryzabase/top/top.jsp). The database contains information about rice development and anatomy, rice mutants, and genetic resources, especially for wild varieties of rice. The anatomical description of rice development is unique and is the first known representation for rice. Developmental and anatomical descriptions include in situ gene expression data serving as stage and tissue markers. The systematic presentation of a large number of rice mutant and mutant trait genes is indispensable, as is description of research in wild strains, core collections, and their detailed characterization. Several genetic, physical, and expression maps with full genome and cDNA sequences are also combined with biological data in Oryzabase. These datasets, when pooled together, could provide a useful tool for gaining greater knowledge about the life cycle of rice, the relationship between phenotype and gene function, and rice genetic diversity. For exchanging community information, Oryzabase publishes the Rice Genetics Newsletter organized by the Rice Genetics Cooperative and provides a mailing service, rice-e-net/rice-net.

Chromosome Mapping↗

Genome SEGE: a database for 'intronless' genes in eukaryotic genomes.

BACKGROUND: A number of completely sequenced eukaryotic genome data are available in the public domain. Eukaryotic genes are either 'intron containing' or 'intronless'. Eukaryotic 'intronless' genes are interesting datasets for comparative genomics and evolutionary studies. The SEGE database containing a collection of eukaryotic single exon genes is available. However, SEGE is derived using GenBank. The redundant, incomplete and heterogeneous qualities of GenBank data are a bottleneck for biological investigation in comparative genomics and evolutionary studies. Such studies often require representative gene sets from each genome and this is possible only by deriving specific datasets from completely sequenced genome data. Thus Genome SEGE, a database for 'intronless' genes in completely sequenced eukaryotic genomes, has been constructed. AVAILABILITY: http://sege.ntu.edu.sg/wester/intronless DESCRIPTION: Eukaryotic 'intronless' genes are extracted from nine completely sequenced genomes (four of which are unicellular and five of which are multi-cellular). The complete dataset is available for download. Data subsets are also available for 'intronless' pseudo-genes. The database provides information on the distribution of 'intronless' genes in different genomes together with their length distributions in each genome. Additionally, the search tool provides pre-computed PROSITE motifs for each sequence in the database with appropriate hyperlinks to InterPro. A search facility is also available through the web server. CONCLUSIONS: The unique features that distinguish Genome SEGE from SEGE is the service providing representative 'intronless' datasets for completely sequenced genomes. 'Intronless' gene sets available in this database will be of use for subsequent bio-computational analysis in comparative genomics and evolutionary studies. Such analysis may help to revisit the original genome data for re-examination and re-annotation.

Databases, Genetic↗

The planktonic microbiome of the Great Barrier Reef.

Large genome databases have markedly improved our understanding of marine microorganisms1-5. Although these resources have focused on prokaryotes, genomes from many dominant marine lineages, such as Pelagibacter and Prochlorococcus, are conspicuously underrepresented. Here we present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), comprising 5,283 prokaryotic genomes obtained from Great Barrier Reef seawater samples using Nanopore and Illumina sequencing, including a collection of high-quality genomes of underrepresented groups. We show that standard short-read assemblies miss these populations owing to a combination of strain heterogeneity and low-GC-percentage sequencing bias. The GBR-MGD also comprises 20 chromosome-level picoeukaryote and 808,585 viral genomes, including a newly described clade of marine Crassvirales. We demonstrate the utility of the GBR-MGD to identify indicator taxa that can reliably predict the effects of reef management practices, such as the establishment of marine protected zones.

Bacteria↗

Comparative genomic mapping of the bovine Fragile Histidine Triad (FHIT) tumour suppressor gene: characterization of a 2 Mb BAC contig covering the locus, complete annotation of the gene, analysis of cDNA and of physiological expression profiles.

BACKGROUND: The Fragile Histidine Triad gene (FHIT) is an oncosuppressor implicated in many human cancers, including vesical tumors. FHIT is frequently hit by deletions caused by fragility at FRA3B, the most active of human common fragile sites, where FHIT lays. Vesical tumors affect also cattle, including animals grazing in the wild on bracken fern; compounds released by the fern are known to induce chromosome fragility and may trigger cancer with the interplay of latent Papilloma virus. RESULTS: The bovine FHIT was characterized by assembling a contig of 78 BACs. Sequence tags were designed on human exons and introns and used directly to select bovine BACs, or compared with sequence data in the bovine genome database or in the trace archive of the bovine genome sequencing project, and adapted before use. FHIT is split in ten exons like in man, with exons 5 to 9 coding for a 149 amino acids protein. VISTA global alignments between bovine genomic contigs retrieved from the bovine genome database and the human FHIT region were performed. Conservation was extremely high over a 2 Mb region spanning the whole FHIT locus, including the size of introns. Thus, the bovine FHIT covers about 1.6 Mb compared to 1.5 Mb in man. Expression was analyzed by RT-PCR and Northern blot, and was found to be ubiquitous. Four cDNA isoforms were isolated and sequenced, that originate from an alternative usage of three variants of exon 4, revealing a size very close to the major human FHIT cDNAs. CONCLUSION: A comparative genomic approach allowed to assemble a contig of 78 BACs and to completely annotate a 1.6 Mb region spanning the bovine FHIT gene. The findings confirmed the very high level of conservation between human and bovine genomes and the importance of comparative mapping to speed the annotation process of the recently sequenced bovine genome. The detailed knowledge of the genomic FHIT region will allow to study the role of FHIT in bovine cancerogenesis, especially of vesical papillomavirus-associated cancers of the urinary bladder, and will be the basis to define the molecular structure of the bovine homologue of FRA3B, the major common fragile site of the human genome.

Acid Anhydride Hydrolases↗

Natural sequence code representations for compression and rapid searching of human-genome style databases.

Numeric descriptions ('bio-informatic descriptions') of amino acid residues have been developed which will be of value whenever the quality and quantity of information in very large (i.e. 'human genome style') gene and protein sequences is to be compared or manipulated. These codes are as natural as possible by our criteria (the same principles could be used in revision of the criteria). In particular, in storing and searching large amounts of sequence data, natural codes--which relate to the properties of amino acids--can be combined with existing fast-search algorithms but introduce several advantages. The code can be assigned such that sub-selection of bits leads to compressed databases with residues defined less specifically, by classes of properties. The most compressed representation leads to the specification of a residue as polar or non-polar, while the most extended representation used at present also allows specification of, for example, glyco-asparagine and phosphoserine. Preliminary studies on both a supercomputer and smaller machines suggest a 'worst-case' speeding of approximately 4.5-fold. For more intelligent searching, coding extensions mixed with the basic sequence data give the sequence data some of the character of a computer program.

Amino Acid Sequence↗

Pristionchus.org: a genome-centric database of the nematode satellite species Pristionchus pacificus.

Comparative studies have been of invaluable importance to the understanding of evolutionary biology. The evolution of developmental programs can be studied in nematodes at a single cell resolution given their fixed cell lineage. We have established Pristionchus pacificus as a major satellite organism for evolutionary developmental biology relative to Caenorhabditis elegans, the model nematode. Online genomic information to support studies in this satellite system can be accessed at http://www.pristionchus.org. Our web resource offers diverse content covering genome browsing, genetic and physical maps, similarity searches, a community platform and assembly details. Content will be continuously improved as we annotate the P.pacificus genome, and will be an indispensable resource for P.pacificus genomics.

Animals↗

A model of random mass-matching and its use for automated significance testing in mass spectrometric proteome analysis.

A rapid and accurate method for testing the significance of protein identities determined by mass spectrometric analysis of protein digests and genome database searching is presented. The method is based on direct computation using a statistical model of the random matching of measured and theoretical proteolytic peptide masses. Protein identification algorithms typically rank the proteins of a genome database according to a score based on the number of matches between the masses obtained by mass spectrometry analysis and the theoretical proteolytic peptide masses of a database protein. The random matching of experimental and theoretical masses can cause false results. A result is significant only if the score characterizing the result deviates significantly from the score expected from a false result. A distribution of the score (number of matches) for random (false) results is computed directly from our model of the random matching, which allows significance testing under any experimental and database search constraints. In order to mimic protein identification data quality in large-scale proteome projects, low-to-high quality proteolytic peptide mass data were generated in silico and subsequently submitted to a database search program designed to include significance testing based on direct computation. This simulation procedure demonstrates the usefulness of direct significance testing for automatically screening for samples that must be subjected to peptide sequence analysis by e.g. tandem mass spectrometry in order to determine the protein identity.

Algorithms↗

The Indian Genome Variation database (IGVdb): a project overview.

Indian population, comprising of more than a billion people, consists of 4693 communities with several thousands of endogamous groups, 325 functioning languages and 25 scripts. To address the questions related to ethnic diversity, migrations, founder populations, predisposition to complex disorders or pharmacogenomics, one needs to understand the diversity and relatedness at the genetic level in such a diverse population. In this backdrop, six constituent laboratories of the Council of Scientific and Industrial Research (CSIR), with funding from the Government of India, initiated a network program on predictive medicine using repeats and single nucleotide polymorphisms. The Indian Genome Variation (IGV) consortium aims to provide data on validated SNPs and repeats, both novel and reported, along with gene duplications, in over a thousand genes, in 15,000 individuals drawn from Indian subpopulations. These genes have been selected on the basis of their relevance as functional and positional candidates in many common diseases including genes relevant to pharmacogenomics. This is the first large-scale comprehensive study of the structure of the Indian population with wide-reaching implications. A comprehensive platform for Indian Genome Variation (IGV) data management, analysis and creation of IGVdb portal has also been developed. The samples are being collected following ethical guidelines of Indian Council of Medical Research (ICMR) and Department of Biotechnology (DBT), India. This paper reveals the structure of the IGV project highlighting its various aspects like genesis, objectives, strategies for selection of genes, identification of the Indian subpopulations, collection of samples and discovery and validation of genetic markers, data analysis and monitoring as well as the project's data release policy.

Databases, Genetic↗