Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

TargetDB: a target registration database for structural genomics projects.

UNLABELLED: TargetDB is a centralized target registration database that includes protein target data from the NIH structural genomics centers and a number of international sites. TargetDB, which is hosted by the Protein Data Bank (RCSB PDB), provides status information on target sequences and tracks their progress through the various stages of protein production and structure determination. A simple search form permits queries based on contributing site, target ID, protein name, sequence, status and other data. The progress of individual targets or entire structural genomics projects may be tracked over time, and target data from all contributing centers may also be downloaded in the XML format. AVAILABILITY: TargetDB is available at http://targetdb.pdb.org/

Amino Acid Sequence↗

Molecular cloning and characterization of a human cDNA and gene encoding a novel acid ceramidase-like protein.

Computer-assisted database analysis of sequences homologous to human acid ceramidase (ASAH) revealed a 1233-bp cDNA (previously designated cPj-LTR) whose 266-amino-acid open reading frame had approximately 36% identity with the ASAH polypeptide. Based on this high degree of homology, we undertook further molecular characterization of cPj-LTR and now report the full-length cDNA sequence, complete gene structure (renamed human ASAHL since it is a human acid ceramidase-like sequence), chromosomal location, primer extension and promoter analysis, and transient expression results. The full-length human ASAHL cDNA was 1825 bp and contained an open-reading frame encoding a 359-amino-acid polypeptide that was 33% identical and 69% similar to the ASAH polypeptide over its entire length. Numerous short regions of complete identity were observed between these two sequences and two sequences obtained from the Caenorhabditis elegans genome database. The 30-kb human ASAHL genomic sequence contained 11 exons, which ranged in size from 26 to 671 bp, and 10 introns, which ranged from 150 bp to 6.4 kb. The gene was localized to the chromosomal region 4q21.1 by fluorescence in situ hybridization analysis. Northern blotting experiments revealed a major 2.0-kb ASAHL transcript that was expressed at high levels in the liver and kidney, but at relatively low levels in other tissues such as the lung, heart, and brain. Sequence analysis of the 5'-flanking region of the human ASAHL gene revealed a putative promoter region that lacked a TATA box and was GC rich, typical features of a housekeeping gene promoter, as well as several tissue-specific and/or hormone-induced transcription regulatory sites. 5'-Deletion analysis localized the promoter activity to a 1. 1-kb fragment within this region. A major transcription start site also was located 72 bp upstream from the ATG translation initiation site by primer extension analysis. Expression analysis of a green fluorescence protein/ASAHL fusion protein in COS-1 cells revealed a punctate, perinuclear distribution, although no acid ceramidase activity was detected in the transfected cells using a fluorescence-based in vitro assay system.

5' Untranslated Regions↗

Phylogenetic detection of conserved gene clusters in microbial genomes.

BACKGROUND: Microbial genomes contain an abundance of genes with conserved proximity forming clusters on the chromosome. However, the conservation can be a result of many factors such as vertical inheritance, or functional selection. Thus, identification of conserved gene clusters that are under functional selection provides an effective channel for gene annotation, microarray screening, and pathway reconstruction. The problem of devising a robust method to identify these conserved gene clusters and to evaluate the significance of the conservation in multiple genomes has a number of implications for comparative, evolutionary and functional genomics as well as synthetic biology. RESULTS: In this paper we describe a new method for detecting conserved gene clusters that incorporates the information captured by a genome phylogenetic tree. We show that our method can overcome the common problem of overestimation of significance due to the bias in the genome database and thereby achieve better accuracy when detecting functionally connected gene clusters. Our results can be accessed at database GeneChords http://genomics10.bu.edu/GeneChords. CONCLUSION: The methodology described in this paper gives a scalable framework for discovering conserved gene clusters in microbial genomes. It serves as a platform for many other functional genomic analyses in microorganisms, such as operon prediction, regulatory site prediction, functional annotation of genes, evolutionary origin and development of gene clusters.

Algorithms↗

Sequence analysis of DNA randomly amplified from the Saccharomyces cerevisiae genome.

Despite its widespread use, the molecular basis of random amplification is poorly understood. Here the basis of random amplification has been investigated by cloning and sequencing the products of a random amplification of polymorphic DNA (RAPD) amplification from Saccharomyces cerevisiae DNA. The genomic origin of the amplified products was determined by sequence comparison with the S. cerevisiae Genome Database (SGD). This allowed analysis of the degree of identity between the random primer and the primer binding sites on the genome. There was no relationship between RAPD size, GC content and relative abundance. The degree of matching between the primer and the primer binding sites increased towards the 3; end of the primer and decreased towards the 5; end. The maximum number of mismatches observed between primer and primer binding sites was never more than one between positions 1-7 of the primer. Nucleotide compositional biases were also observed upstream and downstream of the primer binding site with a marked preference for AT richness upstream of the primer binding sites and for a GC preference directly following the 3; end of the primer. These findings have important ramifications for primer design for multiplex, low stringency and degenerate polymerase chain reaction (PCR).

Base Sequence↗

Targeting enzymes involved in spermidine metabolism of parasitic protozoa--a possible new strategy for anti-parasitic treatment.

Sequencing data obtained from the Plasmodium, Anopheles gambiae and human genome projects provide a new basis for drug and vaccine development. One of the most characteristic features in the process of drug development against parasitic protozoa is target identification in a biological pathway. The next step must be a structure-based rational drug design if the target is not only present in the parasite. In mouse models of malaria, such drugs should be tested for efficacy of the new therapies. Here, we present data that pinpoint the existence of two enzymes of the polyamine pathway involved in spermidine metabolism in P. falciparum, i.e. deoxyhypusine synthase (DHS; EC 1.1.1.249) and homospermidine synthase (HSS; EC 2.5.1.45). Recent data obtained from the malaria genome databases showed that at least a putative gene encoding DHS is present in the parasite. Sequencing data from the P. falciparum genome project prove that the eukaryotic initiation factor eIF5A (the substrate for DHS) exists in P. falciparum. Here, we present the amino acid sequence of eIF5A from P. vivax, which causes tertiary malaria. EIF5A from P. vivax shows 82% nucleic acid and 97% amino acid identity to its homologue from P. falciparum. GC/MS data and inhibitor studies with agmatine prove that the triamine homospermidine occurs in the parasite. These data suggest a separate locus encoding HSS in P. falciparum. The hss gene recruits from the dhs gene in eukaryotes. Here, we present genomic DNA fragments obtained by amplification with primers of a conserved region (amino acid positions 550-1,043) between the putative P. falciparum DHS gene ( dhs) and the HSS gene ( hss) from the plant Senecio vulgaris (Asteraceae). The amplification product from different P. falciparum strains reveals differences in sequence identity, compared with the putative dhs gene from P. falciparum strain 3D7. Expression of the full-length clone and determination of HSS-specific activity will finally prove whether a separate region encoding HSS exists.

Alkyl and Aryl Transferases↗

Reevaluation of the predicted gene structure of Dictyostelium cystatin A3 (cpiC) by nucleotide sequence determination of its cDNA* and its phylogenetic position in the cystatin superfamily.

Cystatins, cysteine protease inhibitors, are widely distributed among eukaryotes. We reevaluated the structure of the gene cpiC, a gene encoding the third identified member of cystatin family (cystatin A3) that was predicted in the genome database of the social amoeba Dictyostelium discoidium (dictyBase) but remained controversial. We determined the sequences of cDNA and PCR-amplified genomic DNA fragment and found a critical error in the registered nucleotide sequence. The corrected cystatin A3 gene has an open reading frame (ORF) without intron sequence interruption and encodes 94 amino acids (aa), in contrast to the previously predicted sequence of either 80, 82 or 118 aa. The cDNA has an unusual internal poly(A) sequence of 31 adenines, which immediately follows the translation termination codon (TAA) located 146 nucleotides upstream of the post-transcriptional polyadenylation site. The amino acid sequence of Dictyostelium cystatin A3 shows a high similarity to those of previously reported Dictyostelium cystatins as well as Family I cystatins of higher eukaryotes.

Amino Acid Sequence↗

Physiological and molecular interaction in the host-parasitoid system Heliothis virescens-Toxoneuron nigriceps: current status and future perspectives.

Toxoneuron nigriceps (Viereck) (Hymenoptera, Braconidae) is an endophagous parasitoid of the tobacco budworm Heliothis virescens (F.) (Lepidoptera, Noctuidae). Parasitized H. virescens larvae are developmentally arrested and show a complex array of pathological symptoms ranging from the suppression of the immune response to an alteration of ecdysone biosynthesis and metabolism. Most of these pathological syndromes are induced by the polydnavirus associated with T. nigriceps (TnBV). An overview of our recent research work on this system is described herein. The mechanisms involved in the disruption of the host hormonal balance have been further investigated, allowing to better define the physiological model previously proposed. A functional genomic approach has been undertaken to identify TnBV genes expressed in the host and to assess their role in the major parasitoid-induced pathologies. Some TnBV genes cloned so far are novel and do not show any similarity with genes already available in genomic databases, while others code for proteins having conserved domains, such as aspartic proteases and tyrosine phosphatases. Sequencing of the entire TnBV genome is in progress and will considerably contribute to the understanding of the molecular bases of parasitoid-induced host alterations.

Amino Acid Sequence↗

Quantitative traits in plants: beyond the QTL.

Phenotypic variation for quantitative traits results from segregation at multiple quantitative trait loci (QTL), the effects of which are modified by the internal and external environments. Because of their favorable genetic attributes (e.g. short generation time, large families and tolerance to inbreeding), plants are often used to test new concepts in quantitative trait analysis. Thus far, the molecular basis underlying allelic variation at QTL is similar to the identified variation for simple mendelian loci; namely, alterations in gene expression or protein function. Further comprehensive dissection of complex phenotypes will depend on our ability to link genetic components of the QTL variation to genomic databases.

Alleles↗

Simple sequence repeats in the Helicobacter pylori genome.

We describe an integrated system for the analysis of DNA sequence motifs within complete bacterial genome sequences. This system is based around ACeDB, a genome database with an integrated graphical user interface; we identify and display motifs in the context of genetic, sequence and bibliographic data. Tomb et aL (1997) previously reported the identification of contingency genes in Helicobacter pylori through their association with homopolymeric tracts and dinucleotide repeats. With this as a starting point, we validated the system by a search for this type of repeat and used the contextual information to assess the likelihood that they mediate phase variation in the associated open reading frames (ORFs). We found all of the repeats previously described, and identified 27 putative phase-variable genes (including 17 previously described). These could be divided into three groups: lipopolysaccharide (LPS) biosynthesis, cell-surface-associated proteins and DNA restriction/modification systems. Five of the putative genes did not have obvious homologues in any of the public domain sequence databases. The reading frame of some ORFs was disrupted by the presence of the repeats, including the alpha(1-2) fucosyltransferase gene, necessary for the synthesis of the Lewis Y epitope. An additional benefit of this approach is that the results of each search can be analysed further and compared with those from other genomes. This revealed that H. pylori has an unusually high frequency of homopurine:homopyrimidine repeats suggesting mechanistic biases that favour their presence and instability.

Base Sequence↗

The Medicago Genome Initiative: a model legume database.

The Medicago Genome Initiative (MGI) is a database of EST sequences of the model legume MEDICAGO: truncatula. The database is available to the public and has resulted from a collaborative research effort between the Samuel Roberts Noble Foundation and the National Center for Genome Resources to investigate the genome of M.truncatula. MGI is part of the greater integrated MEDICAGO: functional genomics program at the Noble Foundation (http://www.noble.org ), which is taking a global approach in studying the genetic and biochemical events associated with the growth, development and environmental interactions of this model legume. Our approach will include: large-scale EST sequencing, gene expression profiling, the generation of M.truncatula activation-tagged and promoter trap insertion mutants, high-throughput metabolic profiling, and proteome studies. These multidisciplinary information pools will be interfaced with one another to provide scientists with an integrated, holistic set of tools to address fundamental questions pertaining to legume biology. The public interface to the MGI database can be accessed at http://www.ncgr.org/research/mgi.

Computational Biology↗

Ontologies for molecular biology.

Molecular biology has a communication problem. There are many databases using their own labels and categories for storing data objects and some using identical labels and categories but with a different meaning. A prominent example is the concept "gene" which is used with different semantics by major international genomic databases. Ontologies are one means to provide a semantic repository to systematically order relevant concepts in molecular biology and to bridge the different notions in various databases by explicitly specifying the meaning of and relation between the fundamental concepts in an application domain. Here, the upper level and a database branch of a prospective ontology for molecular biology (OMB) is presented and compared to other ontologies with respect to suitability for molecular biology (http:/(/)igd.rz-berlin.mpg.de/approximately www/oe/mbo.html).

Chromosome Mapping↗

GeneDB: a resource for prokaryotic and eukaryotic organisms.

GeneDB (http://www.genedb.org/) is a genome database for prokaryotic and eukaryotic organisms. The resource provides a portal through which data generated by the Pathogen Sequencing Unit at the Wellcome Trust Sanger Institute and other collaborating sequencing centres can be made publicly available. It combines data from finished and ongoing genome and expressed sequence tag (EST) projects with curated annotation, that can be searched, sorted and downloaded, using a single web based resource. The current release stores 11 datasets of which six are curated and maintained by biologists, who review and incorporate information from the scientific literature, public databases and the respective research communities.

Animals↗

Genomic organization and sequence variation of the human integrin subunit alpha8 gene (ITGA8).

The integrin alpha8 is highly expressed during kidney and lung development. alpha8-deficient mice display abnormal renal development suggesting that alpha8 plays a critical role in organogenesis. Therefore, it would be of considerable interest to understand the genomic structure, localization and sequence variation of the alpha8 gene. Using FISH and genomic database analysis, we show that alpha8 gene maps to chromosome 10p13 and consists of >200 kbp organized into 30 exons. Examination of 47 individuals from two different ethnic groups (European and African descent) identified 286 varying sites. The diversity of alpha8 is comparable to that of other regions within the human genome. Eight of the varying sites were located in the coding regions: six resulted in nonsynonymous substitutions of which two lead to non-conservative changes in protein. None of the sites showed significant deviation from Hardy-Weinberg equilibrium. We mapped the coding region single nucleotide polymorphisms (SNPs) onto a model of the predicted alpha8 structure and found all the SNPs were located in the "calf" of the extracellular domain. In the European population, the linkage disequilibrium statistic D' showed three blocks of relatively non-recombinant regions in the alpha8 gene while the African population showed more evidence of recombination. The observed patterns of the linkage disequilibrium statistic R2 suggest that a large number of sites will need to be genotyped to ensure coverage of the entire gene for genetic association studies. Identification of the sequence variation will allow genetic association studies of alpha8 in kidney and lung disease.

Base Sequence↗

FRA1E common fragile site breaks map within a 370kilobase pair region and disrupt the dihydropyrimidine dehydrogenase gene (DPYD).

Common fragile sites represent components of normal chromosome structure that are particularly prone to breakage under replication stress. Although the cytogenetic locations of 88 common fragile sites are listed in the Genome database, the DNA at only 14 of them has been defined and characterized at the molecular level. Here, we identify the precise genomic position of the common fragile site FRA1E, mapped to the chromosomal band 1p21.2, and characterize the genetic complexity of the fragile DNA sequence. We show that FRA1E extends over 370kb within the dihydropyrimidine dehydrogenase (DPYD) gene, which genomically spans approximately 840kb. The 185kb region of the highest fragility, which accounts for 86% of all observed breaks at FRA1E, encompasses the central part of DPYD including exons 13-16. DPYD encodes dihydropyrimidine dehydrogenase (DPD), which is the first and rate-limiting enzyme in a three-step metabolic pathway involved in degradation of the pyrimidine bases uracil and thymine. Deficiency in human DPD is associated with autosomal recessive disease, thymine-uraciluria, and with severe 5-fluorouracil toxicity in cancer patients. To which extent the disruption of the DPYD gene by the fragile site break is only transient, followed by DNA repair to restore the original structure, or occasionally may result in genomic damage associated with human disease remains to be determined.

Aphidicolin↗

Identification of the Vibrio cholerae type 4 prepilin peptidase required for cholera toxin secretion and pilus formation.

Cholera toxin secretion is dependent upon the extracellular protein secretion apparatus encoded by the eps gene locus of Vibrio cholerae. Although the eps gene locus encodes several type four prepilin-like proteins, the peptidase responsible for processing these proteins has not been identified. This report describes the identification of a prepilin peptidase from the V. cholerae genomic database by virtue of its homology with the PilD prepilin peptidase of Pseudomonas aeruginosa. Plasmid disruption or deletion of this peptidase gene in either EI Tor or classical V. cholerae O1 biotype strains results in a dramatic decrease in cholera toxin secretion. In the case of the EI Tor biotype mutants, surface expression of the type 4 pilus responsible for mannose-sensitive haemagglutination is abolished. The cloned V. cholerae peptidase processes either EpsI or MshA preproteins when co-expressed in E. coli. Mutation of the V. cholerae peptidase gene also results in a defect in virulence and decreased levels of OmpU. The V. cholerae peptidase gene sequence shows 80% homology with the Vibrio vulnificus VvpD type 4 prepilin peptidase required for pilus assembly and cytolysin secretion in V. vulnificus. Accordingly, the V. cholerae type 4 prepilin peptidase required for pilus assembly and cholera toxin secretion has been designated VcpD.

Adhesins, Bacterial↗

HOWDY: an integrated database system for human genome research.

HOWDY is an integrated database system for accessing and analyzing human genomic information (http://www-alis.tokyo.jst.go.jp/HOWDY/). HOWDY stores information about relationships between genetic objects and the data extracted from a number of databases. HOWDY consists of an Internet accessible user interface that allows thorough searching of the human genomic databases using the gene symbols and their aliases. It also permits flexible editing of the sequence data. The database can be searched using simple words and the search can be restricted to a specific cytogenetic location. Linear maps displaying markers and genes on contig sequences are available, from which an object can be chosen. Any search starting point identifies all the information matching the query. HOWDY provides a convenient search environment of human genomic data for scientists unsure which database is most appropriate for their search.

Chromosome Mapping↗

The structure of the FMRFamide receptor and activity of the cardioexcitatory neuropeptide are conserved in mosquito.

Numerous peptides are structurally related to the cardioexcitatory tetrapeptide FMRFamide. One subgroup of FMRFamide-related peptides (FaRPs) contains an FMRFamide C terminus. Searches of the Drosophila melanogaster genome database identified the first invertebrate FMRFamide G-protein coupled receptor (GPCR), DrmFMRFa-R (Cazzamali and Grimmelikhuijzen, Meeusen et al., 2002). In order to explore molecular mechanisms involved in FMRFamide signal transduction we identified a receptor from the malaria mosquito Anopheles gambiae genome (Holt et al., 2002), AngFMRFa-R, and compared its structure to DrmFMRFa-R. The cytoplasmic loops, extracellular loops, and transmembrane regions are highly conserved between these two FMRFamide receptors. Another subgroup of FaRPs is the sulfakinins which are represented by the consensus structure -XDYGHMRFamide, where X is D or E (Nichols, 2003). We compared AngFMRFa-R and DrmFMRFa-R to the A. gambiae sulfakinin receptors, ASK-R1 and ASK-R2 ( Duttlinger et al., 2003), and the D. melanogaster sulfakinin receptors, DSK-R1 and DSK-R2 Brody and Cravchik, 2000; Hewes and Taghert, 2001 ). The cytoplasmic loops, extracellular loops, and the transmembrane regions are not highly conserved between the FMRFamide and sulfakinin receptors. In order to explore the role of FMRFamide in mosquito biology we measured the effect of the tetrapeptide on in vivo heart rate. The tetrapeptide increased the frequency of spontaneous contractions of the larval mosquito heart and, thus, increased heart rate. These data support the conclusion that the structure of the FMRFamide receptor and activity of the cardioexcitatory FMRFamide neuropeptide are conserved in mosquito.

Aedes↗

Knowledge-based selection of targets for structural genomics.

The problem of rational target selection for protein structure determination in structural genomics projects on microbes is addressed. A flexible computational procedure is described that directly incorporates the whole body of annotation available in the PEDANT genome database into the sequence clustering and selection process in order to identify proteins that are likely to possess currently unknown structural domains. Filtering out gene products based on predicted structural features, such as known three-dimensional structures and transmembrane regions, allows one to reduce the complexity of neighbor relationships between sequences and all but eliminates the need for further partitioning of single-linkage clusters into disjoint protein groups corresponding to homologous families. The results of a large-scale computation experiment in which exemplary target selection for 32 prokaryotic genomes was conducted are presented.

Algorithms↗