Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Evolution of vertebrate E protein transcription factors: comparative analysis of the E protein gene family in Takifugu rubripes and humans.

E proteins are essential for B lymphocyte development and function, including immunoglobulin (Ig) gene rearrangement and expression. Previous studies of B cells in the channel catfish (Ictalurus punctatus) identified E protein homologs that are capable of binding the muE5 motif and driving a strong transcriptional response. There are three E protein genes in mammals, HEB (TCF12), E2A (TCF3), and E2-2 (TCF4). The major expressed E proteins found in catfish B cells are homologs of HEB and of E2A. Here we sought to define the complete family of E protein genes in a teleost fish, Takifugu rubripes, taking advantage of the completed genome sequence. The catfish CFEB (HEB homolog) sequence identified homologous E-protein-encoding sequences in five scaffolds in the Takifugu genome database. Detailed comparative analysis with the human genome revealed the presence of five E protein homologs in Takifugu. Single genes orthologous to HEB and to E2-2 were identified. In contrast, two members of the E2A gene family were identified in Takifugu; one of these shows the alternative processing of transcripts that identifies it as the ortholog of the E12/E47-encoding mammalian E2A gene, whereas the second Takifugu E2A gene has no predicted alternative splice products. A novel fifth E protein gene (EX) was identified in Takifugu. Phylogenetic analysis revealed four E protein branches among vertebrates: EX, E2A, HEB, and E2-2.

Animals↗

Genomic regulation of natural variation in cortical and noncortical brain volume.

BACKGROUND: The relative growth of the neocortex parallels the emergence of complex cognitive functions across species. To determine the regions of the mammalian genome responsible for natural variations in cortical volume, we conducted a complex trait analysis using 34 strains of recombinant inbred (Rl) strains of mice (BXD), as well as their two parental strains (C57BL/6J and DBA/2J). We measured both neocortical volume and total brain volume in 155 coronally sectioned mouse brains that were Nissl stained and embedded in celloidin. After correction for shrinkage, the measured cortical and noncortical brain volumes were entered into a multiple regression analysis, which removed the effects of body size and age from the measurements. Marker regression and interval mapping were computed using WebQTL. RESULTS: An ANOVA revealed that more than half of the variance of these regressed phenotypes is genetically determined. We then identified the regions of the genome regulating this heritability. We located genomic regions in which a linkage disequilibrium was present using WebQTL as both a mapping engine and genomic database. For neocortex, we found a genome-wide significant quantitative trait locus (QTL) on chromosome 11 (marker D11Mit19), as well as a suggestive QTL on chromosome 16 (marker D16Mit100). In contrast, for noncortex the effect of chromosome 11 was markedly reduced, and a significant QTL appeared on chromosome 19 (D19Mit22). CONCLUSION: This classic pattern of double dissociation argues strongly for different genetic factors regulating relative cortical size, as opposed to brain volume more generally. It is likely, however, that the effects of proximal chromosome 11 extend beyond the neocortex strictly defined. An analysis of single nucleotide polymorphisms in these regions indicated that ciliary neurotrophic factor (Cntf) is quite possibly the gene underlying the noncortical QTL. Evidence for a candidate gene modulating neocortical volume was much weaker, but Otx1 deserves further consideration.

Animals↗

Evolving a legacy system: restructuring the Mendelian Inheritance in Man database.

Mendelian Inheritance in Man (MIM) is an encyclopedia of medical genetics that has been in electronic form for over 30 years. In its lifetime, MIM has undergone many organizational and software changes. In 1994, a major transition was made based on three basic principles: industry standards, open systems architecture, and extensibility. The resulting MIM database allows users to navigate to other genomic databases, permits the delivery of multimedia information, and improves the quality of data. The new MIM database also improved its administration because of 1) an internal format that enforces consistency; 2) a lower maintenance cost of software; and 3) a better ability to migrate MIM content. In addition, the new architecture will allow MIM easily to adopt emergent technologies as they mature.

Computer Systems↗

Impact of Human Genome Project on treatment of frail and edentulous patients.

OBJECTIVE: Because of ongoing increases in life expectancy and deferment of edentulousness to older age, dentists are facing a different challenge to satisfy elderly denture wearers with a higher prevalence of chronic diseases. This discussion introduces the Human Genome databases as novel and powerful resources to re-examine the core problems experienced by frail and edentulous patients. BACKGROUND: Recent studies demonstrated that mandibular implant overdentures do not necessarily increase masticatory function, perception and satisfaction in denture wearers with adequate edentulous residual ridges. It has been demonstrated that the rate of edentulous residual ridge resorption significantly varies among individuals. The prognosis and cost-effectiveness of denture treatment, with or without implants, may largely depend on how the edentulous ridge is maintained. However, reliable clinical methods permitting dentists to predict the long-term health of the edentulous residual ridge are lacking. MATERIALS AND METHODS: With the completion of the Human Genome Project, the genomic sequence database from this multinational consortium will provide a unique resource to determine the genetic basis of similarity and diversity of humans. RESULTS: One base pair in every 100 to 300 base pairs of the genome sequence varies among humans, suggesting that genetic diagnosis using the single nucleotide polymorphisms (SNPs) may provide a novel opportunity to differentiate our edentulous patients. CONCLUSIONS: Future dental service for the elderly will require a personalized care paradigm, using highly sensitive diagnostic technology such as SNP genomic analysis, for recommending the treatment with greatest potential benefit.

Aged↗

DIAN: a novel algorithm for genome ontological classification.

Faced with the determination of many completely sequenced genomes, computational biology is now faced with the challenge of interpreting the significance of these data sets. A multiplicity of data-related problems impedes this goal: Biological annotations associated with raw data are often not normalized, and the data themselves are often poorly interrelated and their interpretation unclear. All of these problems make interpretation of genomic databases increasingly difficult. With the current explosion of sequences now available from the human genome as well as from model organisms, the importance of sorting this vast amount of conceptually unstructured source data into a limited universe of genes, proteins, functions, structures, and pathways has become a bottleneck for the field. To address this problem, we have developed a method of interrelating data sources by applying a novel method of associating biological objects to ontologies. We have developed an intelligent knowledge-based algorithm, to support biological knowledge mapping, and, in particular, to facilitate the interpretation of genomic data. In this respect, the method makes it possible to inventory genomes by collapsing multiple types of annotations and normalizing them to various ontologies. By relying on a conceptual view of the genome, researchers can now easily navigate the human genome in a biologically intuitive, scientifically accurate manner.

Algorithms↗

IMGT, the international ImMunoGeneTics information system, http://imgt.cines.fr: the reference in immunoinformatics.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr), is a high quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC) and related proteins of the immune system of human and other vertebrates, created in 1989, by the Laboratoire d'ImmunoGénétique Moléculaire (LIGM), at the Université Montpellier II, CNRS, Montpellier, France. IMGT is the global reference in immunogenetics and immunoinformatics and provides a common access to standardized data which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes three sequence databases (IMGT/LIGM-DB, IMGT/MHC-DB hosted at EBI, IMGT/PRIMER-DB), one genome database (IMGT/GENE-DB), one 3D structure database (IMGT/3Dstructure-DB), Web resources comprising 8000 HTML pages ("IMGT Marie-Paule page") and interactive tools for sequence (IMGT/V-QUEST, IMGT/JunctionAnalysis, IMGT/Allele-Align, IMGT/PhyloGene) and genome (IMGT/GeneSearch, IMGT/GeneView, IMGT/LocusView) analysis. IMGT data are expertly annotated according to the rules of the IMGT Scientific chart, based on the IMGT-ONTOLOGY concepts. IMGT tools are particularly useful for the analysis of the IG and TR repertoires in physiological normal and pathological situations. IMGT has important applications in medical research (repertoire analysis in autoimmune diseases, AIDS, leukemias, lymphomas, myelomas), biotechnology related to antibody engineering (phage displays, combinatorial libraries) and therapeutic approaches (graft, immunotherapy). IMGT is freely available at http://imgt.cines.fr.

Animals↗

An online database for brain disease research.

BACKGROUND: The Stanley Medical Research Institute online genomics database (SMRIDB) is a comprehensive web-based system for understanding the genetic effects of human brain disease (i.e. bipolar, schizophrenia, and depression). This database contains fully annotated clinical metadata and gene expression patterns generated within 12 controlled studies across 6 different microarray platforms. DESCRIPTION: A thorough collection of gene expression summaries are provided, inclusive of patient demographics, disease subclasses, regulated biological pathways, and functional classifications. CONCLUSION: The combination of database content, structure, and query speed offers researchers an efficient tool for data mining of brain disease complete with information such as: cross-platform comparisons, biomarkers elucidation for target discovery, and lifestyle/demographic associations to brain diseases.

Bipolar Disorder↗

Human gene mutation database-a biomedical information and research resource.

Although 20 years have elapsed since the first single basepair substitution underlying an inherited disease in humans was characterised at the DNA level, the initiative has only recently been taken to establish central database resources for pathological genetic variants. Disease-associated gene lesions are currently collected and publicised by the Human Gene Mutation Database (HGMD) in Cardiff, locus-specific mutation databases, and to some extent also by the Genome Database (GDB) and Online Mendelian Inheritance in Man (OMIM). To date, HGMD represents the only comprehensive and publicly available database of gene lesions underlying human inherited disease. By July 1999, HGMD contained over 18,000 different mutations from some 900 human genes, the majority being single basepair substitutions. In addition to its potential as an information resource for clinicians and genetic counsellors, HGMD has allowed molecular geneticists to address a variety of biological questions through meta-analysis of the collated data. HGMD also promises to assist research workers in optimising mutation search strategies for a given gene. A questionnaire sent out to, and answered by, the editors of 20 key journals revealed that human genetics journals are increasingly reluctant to publish mutation reports. Electronic data submission and publication facilities are therefore urgently required. The World Wide Web (WWW) provides an excellent medium within which to combine the centralised management of basic mutation data, including rigorous quality control, with the possibility of publishing additional mutation-related information. In response to these needs, HGMD has both instituted a collaboration with Springer-Verlag GmbH, Heidelberg, to potentiate free online submission and electronic publication of human gene mutation data and developed links with the curators of locus-specific mutation databases.

Databases, Factual↗

Conserved structure and promoter sequence similarity in the mouse and human genes encoding the zinc finger factor BERF-1/BFCOL1/ZBP-89.

We have characterized the genomic structure of the mouse Zfp148 gene encoding Beta-Enolase Repressor Factor-1 (BERF-1), a Kruppel-like zinc finger protein involved in the transcriptional regulation of several genes, which is also termed ZBP-89, BFCOL1. The cloned Zfp148 gene spans 110 kb of genomic DNA encompassing the 5'-end region, 9 exons, 8 introns, and the 3'-untranslated region. The promoter region displays the typical features of a housekeeping gene: a high G+C content and the absence of canonical TATA and CAAT boxes consistent with the multiple transcription initiation sites determined by primary extension analysis. Computer-assisted search in the human genome database allowed us to determine that the same genomic structure with identical intron-exon organization is conserved in the human homologue ZNF 148. Functional analysis of the 5'-flanking sequence of the mouse gene indicated that the region from nucleotide -205 to +144, relative to the major transcription start site, contains cis-regulatory elements that promote basal expression. Such sequences and the overall promoter architecture are highly conserved in the human gene. Furthermore, we show that the complex transcription pattern of the Zfp148 gene might be due to a combination of alternative splicing and differential polyadenylation sites utilization.

5' Untranslated Regions↗

Protein expression, genomic structure, and polymorphisms of oculomedin.

PURPOSE: To elucidate protein expression, genomic structure, and genomic polymorphisms of a novel gene, 'oculomedin', that has been cloned as a mechanical stretch-response gene from human trabecular cells in culture. METHODS: Polyclonal antibody was prepared by immunizing rabbits with a chemically synthesized 15-mer peptide of oculomedin. Protein expression was revealed by Western blot analysis after polyacrylamide gel electrophoresis of extracts of mechanically stretched trabecular cells and control trabecular cells in culture as well as retinal tissue. Protein localization was studied immunohistochemically in the human eye section. Genomic structure was determined by searching the GenBank database. Genomic polymorphisms of the coding region in 163 glaucoma patients and 50 normal subjects were detected by PCR amplification and direct sequencing. RESULTS: Western blot analysis showed that oculomedin protein was expressed only in stretched trabecular cells, not in control trabecular cells in culture. Immunohistochemically, oculomedin protein was localized to the trabecular meshwork, Schlemm's canal endothelium, retinal photoreceptor cells, and corneal and conjunctival epithelium. The oculomedin gene (OCLM) consists of two exons which are located inside an intron of a different gene of unknown function (C1orf27) on chromosome 1q25, near and telomeric to myocilin (MYOC). Two types of heterozygous nucleotide substitutions resulting in amino acid changes were found in two of 75 patients with primary open-angle glaucoma, but not at all in patients with other types of glaucoma or in normal subjects. CONCLUSIONS: Oculomedin may play a role in the function of the trabecular meshwork and also in the development of primary open-angle glaucoma.

Animals↗

Saturated BLAST: an automated multiple intermediate sequence search used to detect distant homology.

MOTIVATION: Two proteins can have a similar 3-dimensional structure and biological function, but have sequences sufficiently different that traditional protein sequence comparison algorithms do not identify their relationship. The desire to identify such relations has led to the development of more sensitive sequence alignment strategies. One such strategy is the Intermediate Sequence Search (ISS), which connects two proteins through one or more intermediate sequences. In its brute-force implementation, ISS is a strategy that repetitively uses the results of the previous query as new search seeds, making it time-consuming and difficult to analyze. RESULTS: Saturated BLAST is a package that performs ISS in an efficient and automated manner. It was developed using Perl and Perl/Tk and implemented on the LINUX operating system. Starting with a protein sequence, Saturated BLAST runs a BLAST search and identifies representative sequences for the next generation of searches. The procedure is run until convergence or until some predefined criteria are met. Saturated BLAST has a friendly graphic user interface, a built-in BLAST result parser, several multiple alignment tools, clustering algorithms and various filters for the elimination of false positives, thereby providing an easy way to edit, visualize, analyze, monitor and control the search. Besides detecting remote homologies, Saturated BLAST can be used to maintain protein family databases and to search for new genes in genomic databases.

Algorithms↗

THoR: a tool for domain discovery and curation of multiple alignments.

We describe a tool, THoR, that automatically creates and curates multiple sequence alignments representing protein domains. This exploits both PSI-BLAST and HMMER algorithms and provides an accurate and comprehensive alignment for any domain family. The entire process is designed for use via a web-browser, with simple links and cross-references to relevant information, to assist the assessment of biological significance. THoR has been benchmarked for accuracy using the SMART and pufferfish genome databases.

Algorithms↗

SIGMA: a system for integrative genomic microarray analysis of cancer genomes.

BACKGROUND: The prevalence of high resolution profiling of genomes has created a need for the integrative analysis of information generated from multiple methodologies and platforms. Although the majority of data in the public domain are gene expression profiles, and expression analysis software are available, the increase of array CGH studies has enabled integration of high throughput genomic and gene expression datasets. However, tools for direct mining and analysis of array CGH data are limited. Hence, there is a great need for analytical and display software tailored to cross platform integrative analysis of cancer genomes. RESULTS: We have created a user-friendly java application to facilitate sophisticated visualization and analysis such as cross-tumor and cross-platform comparisons. To demonstrate the utility of this software, we assembled array CGH data representing Affymetrix SNP chip, Stanford cDNA arrays and whole genome tiling path array platforms for cross comparison. This cancer genome database contains 267 profiles from commonly used cancer cell lines representing 14 different tissue types. CONCLUSION: In this study we have developed an application for the visualization and analysis of data from high resolution array CGH platforms that can be adapted for analysis of multiple types of high throughput genomic datasets. Furthermore, we invite researchers using array CGH technology to deposit both their raw and processed data, as this will be a continually expanding database of cancer genomes. This publicly available resource, the System for Integrative Genomic Microarray Analysis (SIGMA) of cancer genomes, can be accessed at http://sigma.bccrc.ca.

Adenocarcinoma↗

A bioinformatics tool to select sequences for microarray studies of mouse models of oncogenesis.

UNLABELLED: One of the challenges to the effective utilization of cDNA microarray analysis in mouse models of oncogenesis is the choice of a critical set of probes that are informative for human disease. Given the thousands of genes with a potential role in human oncogenesis and the hundreds of thousands of mouse sequences available for use as probes, selection of an informative set of mouse probes can be an overwhelming task. We have developed a web based sequence mining tool using DataBase Independent (DBI) Perl to annotate publicly available sequences. The Mouse Oncochip Design Tool uses the Mouse Genome Database (MGD) developed and maintained by the Jackson Laboratories for mouse DNA sequences. There are over 380 000 sequences in their database. The output list has been ordered to present the genes more likely to be informative in a mouse model of human cancer using a candidate set of oncogenes to order the list. Mouse sequences that represent genes that are homologous with a member of a human oncogene set are listed first. In addition it provides a set of links for information on clone source gene function. CONTACT: http://nciarray.nci.nih.gov/cgi-bin/me/mouse_design.cgi

Animals↗

QIS: A framework for biomedical database federation.

Query Integrator System (QIS) is a database mediator framework intended to address robust data integration from continuously changing heterogeneous data sources in the biosciences. Currently in the advanced prototype stage, it is being used on a production basis to integrate data from neuroscience databases developed for the SenseLab project at Yale University with external neuroscience and genomics databases. The QIS framework uses standard technologies and is intended to be deployable by administrators with a moderate level of technological expertise: It comes with various tools, such as interfaces for the design of distributed queries. The QIS architecture is based on a set of distributed network-based servers, data source servers, integration servers, and ontology servers, that exchange metadata as well as mappings of both metadata and data elements to elements in an ontology. Metadata version difference determination coupled with decomposition of stored queries is used as the basis for partial query recovery when the schema of data sources alters.

Computer Communication Networks↗

Integration of cytogenetic data with genome maps and available probes: present status and future promise.

The National Cancer Institute has established an initiative, called the Cancer Chromosome Aberration Project (Ccap), in order to link and integrate the physical and genetic maps of the human genome with cytogenetic data and the location of chromosomal rearrangements in human diseases. This goal will be achieved by high-resolution fluorescence in situ hybridization (FISH) mapping of colony-purified bacterial artificial chromosome (BAC) clones spaced at 1-to 2-Mb intervals across the entire genome. All BAC clones will be anchored on the physical map by the presence of a mapped sequence tagged site (STS). The generation of a publicly accessible clone repository will allow convenient distribution of these BACs. Ccap data can be correlated with other cancer-associated and genomic databases, such as the catalog of chromosomal aberrations in cancer and the emerging full genomic sequence. We anticipate that the use of Ccap clones will expedite and refine the mapping of chromosomal breakpoints. The eventual set of approximately 3,000 Ccap BACs should facilitate the production of BAC-containing DNA chips for assessing copy number of genomic segments by matrix comparative genomic hybridization. In addition, the repository will provide genome-wide tools for defining chromosomal aberrations in cytological specimens by interphase cytogenetics. The Ccap Web site illustrates goals and progress of this initiative (http://www.ncbi.nlm.nih.gov/CCAP/).

Chromosome Mapping↗

The gene complement for proteolysis in the cyanobacterium Synechocystis sp. PCC 6803 and Arabidopsis thaliana chloroplasts.

A set of 62 genes that encode the entire peptidase complement of Synechocystis sp. PCC 6803 has been identified in the genome database of that cyanobacterium. Sequence comparisons with the Arabidopsis genome uncovered the presumably homologous chloroplast components inherited from their cyanobacterial ancestor. A systematic gene disruption approach was chosen to individually inactivate, by customary transformation strategies, the majority of the cyanobacterial genes encoding peptidase subunits that are related to chloroplast enzymes. This allowed classification of the peptidases that are required for cell viability or are involved in specific stress responses. The comparative analysis between Synechocystis and Arabidopsis chloroplast peptidases showed that: (1) homologous enzymes that arose by gene duplications in cyanobacteria are functionally diverse and frequently do not complement each other, (2) the chloroplast appears to house a number of distinct peptidase polypeptide chains of cyanobacterial origin (49) which is comparable with a cyanobacterial cell (62) and (3) the peptidase complement in plastids results from a combination of the loss of some cyanobacterial peptidases and the gain or diversification of subclasses of peptidases. This reorganization in the pattern of proteolytic enzymes may reflect distinct environmental and physiological changes between prokaryotic and organellar systems.

ATP-Dependent Proteases↗