Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Development of the polymerase chain reaction assay based on the canine genome database for detection of monoclonality in B cell lymphoma.

From the canine genome database and its bioinformatic analysis, we identified conserved sequences within the vast majority of 61 variable segments and 1 joining segment of the immunoglobulin heavy chain (IgH) gene, and designed optimal primers for polymerase chain reaction (PCR) amplification directed at these conserved sequences to evaluate the monoclonality of IgH in canine B cell lymphoma. Using the primers, a PCR-based assay was performed on fine-needle aspiration samples of normal, hyperplasia, and malignant lymph nodes and lymphoma cell lines. All fine-needle aspiration samples of five B cell lymphoma cases and the B cell lymphoma line GL-1 exhibited clonal amplification, whereas no amplification was observed in the samples from normal and hyperplasia lymph nodes, cases of T cell lymphoma, and the T cell lymphoma line CL-1. The primers we designed clearly distinguished malignant B lymphocytes from normal, reactive, and malignant T lymphocytes, indicating a potential utility of the primers for PCR-based routine clinical examination for canine B cell lymphoma.

Animals↗

Saccharomyces Genome Database (SGD) provides tools to identify and analyze sequences from Saccharomyces cerevisiae and related sequences from other organisms.

The Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/), a scientific database of the molecular biology and genetics of the yeast Saccharomyces cerevisiae, has recently developed several new resources that allow the comparison and integration of information on a genome-wide scale, enabling the user not only to find detailed information about individual genes, but also to make connections across groups of genes with common features and across different species. The Fungal Alignment Viewer displays alignments of sequences from multiple fungal genomes, while the Sequence Similarity Query tool displays PSI-BLAST alignments of each S.cerevisiae protein with similar proteins from any species whose sequences are contained in the non-redundant (nr) protein data set at NCBI. The Yeast Biochemical Pathways tool integrates groups of genes by their common roles in metabolism and displays the metabolic pathways in a graphical form. Finally, the Find Chromosomal Features search interface provides a versatile tool for querying multiple types of information in SGD.

Amino Acid Sequence↗

A new approach that allows identification of intron-split peptides from mass spectrometric data in genomic databases.

We present a new approach that allows the identification of intron-split peptides from mass spectrometric data in genomic databases. Our algorithm uses small regions of peptide sequence information which are automatically deduced from de novo amino acid sequence predictions together with the molecular mass information of the precursor ion. The sequence predictions are based on selected collision-induced mass spectrometric fragmentation spectra. Fragments of the predicted amino acid sequence are aligned with each of the six frames of the translated genome and the precursor mass information is used to assemble the corresponding tryptic peptides using the sequence as a matrix. Hereby, intron-split peptides can be gathered and in turn verified by mass spectrometric data interpretation tools such as Sequest.

Algorithms↗

Saccharomyces Genome Database provides tools to survey gene expression and functional analysis data.

Upon the completion of the SACCHAROMYCES: cerevisiae genomic sequence in 1996 [Goffeau,A. et al. (1997) NATURE:, 387, 5], several creative and ambitious projects have been initiated to explore the functions of gene products or gene expression on a genome-wide scale. To help researchers take advantage of these projects, the SACCHAROMYCES: Genome Database (SGD) has created two new tools, Function Junction and Expression Connection. Together, the tools form a central resource for querying multiple large-scale analysis projects for data about individual genes. Function Junction provides information from diverse projects that shed light on the role a gene product plays in the cell, while Expression Connection delivers information produced by the ever-increasing number of microarray projects. WWW access to SGD is available at genome-www.stanford. edu/Saccharomyces/.

Databases, Factual↗

Pseudomonas aeruginosa Genome Database and PseudoCAP: facilitating community-based, continually updated, genome annotation.

Using the Pseudomonas aeruginosa Genome Project as a test case, we have developed a database and submission system to facilitate a community-based approach to continually updated genome annotation (http://www.pseudomonas.com). Researchers submit proposed annotation updates through one of three web-based form options which are then subjected to review, and if accepted, entered into both the database and log file of updates with author acknowledgement. In addition, a coordinator continually reviews literature for suitable updates, as we have found such reviews to be the most efficient. Both the annotations database and updates-log database have Boolean search capability with the ability to sort results and download all data or search results as tab-delimited files. To complement this peer-reviewed genome annotation, we also provide a linked GBrowse view which displays alternate annotations. Additional tools and analyses are also integrated, including PseudoCyc, and knockout mutant information. We propose that this database system, with its focus on facilitating flexible queries of the data and providing access to both peer-reviewed annotations as well as alternate annotation information, may be a suitable model for other genome projects wishing to use a continually updated, community-based annotation approach. The source code is freely available under GNU General Public Licence.

DNA, Bacterial↗

Future vision of the GDB human genome database.

In 1973, scientists assembled at the first Human Gene Mapping Workshop to discuss the 64 human genes mapped at that time. In 1989, the GDB Human Genome Database was created to store information on 1, 700 mapped human genes. Ten years later, as the human genome project closes in on the release of the complete DNA sequence holding as many as 100,000 human genes, GDB is evolving to continue to meet the needs of the scientific community. Well known as a resource for data which has been stringently reviewed as part of the curation process, GDB prepares to continue to provide a compilation of the human genome including maps, map objects, polymorphisms, and mutations. As more sites across the Internet are established to share biological information, it becomes increasingly burdensome for the scientist to collect data from all sources of a particular domain. In an attempt to reduce this burden, GDB continues to load data from large genome centres and accept submissions from researchers around the world. Moreover, GDB looks to provide a mechanism to link gene-related information to the human reference sequence. In doing this, GDB plans to establish federated linkages with "boutique" databases around the world that could contain enormous amounts of valuable information about specific genes or chromosomes.

Computer Simulation↗

A practical guide to orient yourself in the labyrinth of genome databases.

The identification of genes involved in human inherited disorders has been revolutionized by the resources produced by the Human Genome Project. In particular, the generation of >1 000 000 human expressed sequence tags (ESTs) has led to the partial identification of a significant percentage of all human genes. In the next 7 years, we will witness another revolution when sequencing of the human genome is complete. The generation of large amounts of genomic data must be accompanied by parallel efforts to make the information easily accessible. Efforts towards this goal have already started, but retrieval of information from genomic databases still remains an arduous task. With practical examples, we will try to show how the currently available information can be exploited usefully, in particular to identify candidate genes for human diseases.

Chromosome Mapping↗

Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases.

MOTIVATION: The prediction of transcription units (TUs, which are similar to operons) is an important problem that has been tackled using many different approaches. The availability of complete microbial genomes has made genome-wide TU predictions possible. Pathway-genome databases (PGDBs) add metabolic and other organizational (i.e. protein complexes) information to the annotated genome, and are able to capture TU organization information. These characteristics of PGDBs make them a suitable framework for the development and implementation of TU predictors. RESULTS: We implemented a TU predictor that uses only intergenic distance and functional classification of genes to predict TU boundaries, and applied it to EcoCyc, our PGDB of Escherichia coli. To this original predictor, we added information on metabolic pathways, protein complexes and transporters, all readily available in EcoCyc, in order to generate an enhanced predictor. The enhanced predictor correctly predicted 80% of the known E.coli TUs (69% of the known operons), a moderate improvement over the original predictor's performance (75% of TUs and 65% of operons correctly predicted), demonstrating that the extra information available in the PGDB does indeed improve prediction performance. Performance of this E.coli-based predictor on a genome other than that of E.coli was tested on BsubCyc, our computationally generated PGDB for Bacillus subtilis, for which a set of 100 known operons is available. Prediction accuracy decreased substantially (46% of the known operons correctly predicted). This was due in part to missing information in BsubCyc, which prevented full use of the predictor's features. The augmented predictor has been implemented as part of our Pathway Tools software suite, and can be used to populate a PGDB with predicted TUs. AVAILABILITY: The TU predictor is included in version 7.0 of the Pathway Tools software suite. Pathway Tools 7.0 is available free of charge to academic institutions and for a fee to commercial enterprises. It runs on Sun Solaris 8, Linux and Windows. TUs predicted on the Caulobacter crescentus and Mycobacterium tuberculosis (H37Rv) genomes are available in our CauloCyc and MtbrvCyc databases, available at the BioCyc web site (http://biocyc.org). To obtain version 7.0 of Pathway Tools, follow the directions in our web site, http://biocyc.org/download.shtml.

Algorithms↗

Building a genome database using an object-oriented approach.

GOBASE is a relational database that integrates data associated with mitochondria and chloroplasts. The most important data in GOBASE, i. e., molecular sequences and taxonomic information, are obtained from the public sequence data repository at the National Center for Biotechnology Information (NCBI), and are validated by our experts. Maintaining a curated genomic database comes with a towering labor cost, due to the shear volume of available genomic sequences and the plethora of annotation errors and omissions in records retrieved from public repositories. Here we describe our approach to increase automation of the database population process, thereby reducing manual intervention. As a first step, we used Unified Modeling Language (UML) to construct a list of potential errors. Each case was evaluated independently, and an expert solution was devised, and represented as a diagram. Subsequently, the UML diagrams were used as templates for writing object-oriented automation programs in the Java programming language.

Databases, Genetic↗

Escherichia coli K12 genomic database.

We have compiled the genomic nucleic acid sequence data of Escherichia coli K12 available from the existing major data collections and from the literature. The collected data are structured as a database for easy access and analysis. The sequence segments in the database are ordered by genetic map position. Sequence redundancy has been completely removed by combining overlapping sequences; therefore, our sequence data are amenable to statistical analysis. We have specified with a plus or minus (+ or -) on which of the two DNA strands the segment exists. The database currently contains a total of 954,392 bp, which corresponds to about 20% of the entire genome size. The sequence data are available on request.

Base Sequence↗

CandidaDB: a genome database for Candida albicans pathogenomics.

CandidaDB is a database dedicated to the genome of the most prevalent systemic fungal pathogen of humans, Candida albicans. CandidaDB is based on an annotation of the Stanford Genome Technology Center C.albicans genome sequence data by the European Galar Fungail Consortium. CandidaDB Release 2.0 (June 2004) contains information pertaining to Assembly 19 of the genome of C.albicans strain SC5314. The current release contains 6244 annotated entries corresponding to 130 tRNA genes and 5917 protein-coding genes. For these, it provides tentative functional assignments along with numerous pre-run analyses that can assist the researcher in the evaluation of gene function for the purpose of specific or large-scale analysis. CandidaDB is based on GenoList, a generic relational data schema and a World Wide Web interface that has been adapted to the handling of eukaryotic genomes. The interface allows users to browse easily through genome data and retrieve information. CandidaDB also provides more elaborate tools, such as pattern searching, that are tightly connected to the overall browsing system. As the C.albicans genome is diploid and still incompletely assembled, CandidaDB provides tools to browse the genome by individual supercontigs and to examine information about allelic sequences obtained from complementary contigs. CandidaDB is accessible at http://genolist.pasteur.fr/CandidaDB.

Candida albicans↗

Mining the Plasmodium genome database to define organellar function: what does the apicoplast do?

Apicomplexan species constitute a diverse group of parasitic protozoa, which are responsible for a wide range of diseases in many organisms. Despite differences in the diseases they cause, these parasites share an underlying biology, from the genetic controls used to differentiate through the complex parasite life cycle, to the basic biochemical pathways employed for intracellular survival, to the distinctive cell biology necessary for host cell attachment and invasion. Different parasites lend themselves to the study of different aspects of parasite biology: Eimeria for biochemical studies, Toxoplasma for molecular genetic and cell biological investigation, etc. The Plasmodium falciparum Genome Project contributes the first large-scale genomic sequence for an apicomplexan parasite. The Plasmodium Genome Database (http://PlasmoDB.org) has been designed to permit individual investigators to ask their own questions, even prior to formal release of the reference P. falciparum genome sequence. As a case in point, PlasmoDB has been exploited to identify metabolic pathways associated with the apicomplexan plastid, or 'apicoplast' - an essential organelle derived by secondary endosymbiosis of an alga, and retention of the algal plastid.

Animals↗

Global prevalence of hereditary hemorrhagic telangiectasia-associated variants estimated by analysis of large-scale genomic databases.

BACKGROUND: Hereditary hemorrhagic telangiectasia (HHT) is an autosomal dominant disorder with an overwhelming hemorrhagic phenotype. It is mainly caused by variants in the ENG and ACVRL1 genes. HHT prevalence is currently estimated to be 1 in 5000 individuals, but the disease is likely underdiagnosed due to variable clinical presentation, misdiagnosis, and delayed recognition. OBJECTIVES: To estimate the global genetic prevalence of HHT-associated variants in ENG and ACVRL1. METHODS: We analyzed 3 large population-scale genomic databases: gnomAD, All of Us, and Regeneron Genetics Center-Million Exome. We considered known pathogenic and likely pathogenic variants of ENG and ACVRL1 and extended the analysis to potentially pathogenic variants passing the pathogenic criteria established by the guidelines for HHT of the American College of Medical Genetics and Genomics/Association for Molecular Pathology. RESULTS: The genetic prevalence of HHT ranged from 1.753 to 2.555 in 5000 individuals, when considering only pathogenic and likely pathogenic variants, and from 2.874 to 4.327 in 5000 individuals, when also potentially pathogenic variants were considered. CONCLUSION: This study assesses the prevalence of HHT-associated variants in the general population. Our unbiased approach demonstrates that the genetic prevalence of the disease is substantially higher than currently estimated.

Humans↗

Identification of novel HrpXo regulons preceded by two cis-acting elements, a plant-inducible promoter box and a -10 box-like sequence, from the genome database of Xanthomonas oryzae pv. oryzae.

A regulatory protein HrpXo of Xanthomonas oryzae pv. oryzae, the causal agent of bacterial leaf blight of rice, is known to control the expression of hrp genes that encode components of a type III secretion system and of some effector protein genes. In this study, we screened novel HrpXo regulons from the genome database of X. oryzae pv. oryzae, searching for ORFs preceded by two predicted sequence motifs, a plant-inducible promoter box-like sequence and a -10 box-like sequence. Using a gus reporter system, nine of 15 ORF candidates were expressed HrpXo dependently. We also showed by base-substituted mutagenesis that both motifs are essential for the expression of the genes.

Bacterial Proteins↗

Gene expression microarray analysis and genome databases facilitate the characterization of a chromosome 22 derived homogeneously staining region.

Karyotype and fluorescence in situ hybridization (FISH) analyses previously identified a homogeneously staining region (HSR) derived from chromosome 22 in OV90, an epithelial ovarian cancer (EOC) cell line. Affymetrix expression microarrays in combination with the UniGene and Human Genome Browser databases were used to identify the candidate genes comprising the amplicon of the HSR, based on comparison of expression profiles of OV90, EOC cell lines lacking HSRs and primary cultures of normal ovarian surface epithelial (NOSE) cells. A group of probe sets displaying a minimum 3-fold overexpression with a high reliability score (P-call) in OV90 were identified which represented genes that mapped within a 1-2 Mb interval on chromosome 22. A large number of probe sets, some of which represent the same genes, displayed no evidence of overexpression and/or low reliability scores (A-call). An investigation of the probe set sequences with the Affymetrix and Sanger Institute Chromosome 22 Group databases revealed that some of the probe sets displaying discordant results for the same gene were complementary to intronic sequences and/or the antisense strand. Microarray results were validated by RT-PCR. Genomic analysis suggests that the HSR was derived from the amplification of a 1.1 Mb interval defined by the chromosomal map positions of ZNF74 and Hs.372662, at 22q11.21. The deduced amplicon is derived from a complex region of chromosome 22 that harbors low-copy repeats (LCRs). The amplicon contains 18 genes as likely targets for gene amplification. This study illustrates that large-scale expression microarray analysis in combination with genome databases is sufficient for deducing target genes associated with amplicons and stresses the importance of investigating probe set design before engaging in validation studies.

Cell Line, Tumor↗

Strategy for identification of novel glucose transporter family members by using internet-based genomic databases.

BACKGROUND: We previously reported that medullary thyroid carcinomas and pheochromocytomas avidly take up the glucose analog fluoro-deoxyglucose on positron emission tomography but do not express any of the known human facilitative glucose transporters. We therefore hypothesized that a novel glucose transporter is responsible for glucose uptake in these tumors. METHODS: Internet-based Expressed Sequence Tags and high throughput genome sequence databases were screened for novel sequences homologous to the known glucose transporters. Derived clones were used to screen cDNA libraries. Sequence comparison and hydropathic analysis of the putative proteins were performed. RESULTS: We identified 2 novel genes (GLUT8 and GLUT9) that are members of the facilitative glucose transporter family. The putative GLUT8 and GLUT9 proteins have 44% and 31% sequence identity to GLUT5 and GLUT3, respectively. Hydropathic analysis showed both have exofacial and transmembrane domains consistent with a hexose transporter. CONCLUSIONS: By using the Expressed Sequence Tags database, we identified novel members of the glucose transporter family. Further work will establish function and expression patterns in medullary thyroid carcinomas and pheochromocytomas. Internet-based genomic databases allow rapid screening and identification of candidate sequences of novel members of human gene families.

Amino Acid Sequence↗