Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Mutation rate variation at human dinucleotide microsatellites.

Mutation is the ultimate source of genetic variation, and mutation rate is thus an important parameter governing the extent of genetic variation. Microsatellites are highly informative genetic markers that have been widely used in genetic studies. While previous studies showed that the mutation rate differs in di-, tri-, and tetranucleotide repeats, how mutation rate distributes within each class of repeat is poorly understood. This study first revealed the pattern of the mutation rate variation within the dinucleotide repeats. Two data sets were used. The first is the allele frequency data from 115 microsatellites with dinucleotide repeats distributed along the human genome in 10 worldwide populations. The second data set is much larger, consisting of the allele frequency of 5252 dinucleotide repeats from the Genome Database. Mutation rate for each locus is estimated through a new homozygosity-based estimator, which has been shown to be unbiased and highly efficient and is reasonably robust against deviations from the single-step model. The mutation rates among loci can be approximated well by a gamma distribution and its shape parameter can be accurately estimated with this approach. This result provides the basic guidelines for analyzing the large-scale genomic data from microsatellite loci.

Alleles↗

Genomic organization and mapping of the gene encoding the PP2A B56gamma regulatory subunit.

Protein phosphatase 2A (PP2A) is a major serine/threonine phosphatase that regulates a wide variety of cellular processes. The enzymatic activity and intracellular localization of PP2A are determined by three distinct families of cellular regulatory subunits (B, B'', and B''). The B' subunit, also known as B56, is the most diverse, consisting of five isoforms (alpha, beta, gamma, delta, and epsilon). The gene encoding B56gamma has been designated as PPP2R5C and encodes three differentially spliced variants: B56gamma1, -gamma2, and -gamma3. However, conflicting chromosomal loci have been reported in human genomic databases. The original cytogenetic mapping placed the gene on chromosome 3p21.3, whereas subsequent studies using radiation hybrid analysis localized PPP2R5C to chromosome 14q. In this study, by radiation hybrid mapping, FISH analysis, BAC clone sequencing, and RT-PCR analysis, we show that the functional gene PPP2R5C exists at 14q32.2 and gives rise to three splicing variants, B56gamma1, -gamma2, and -gamma3, whereas a nonfunctional B56gamma1 pseudogene, PPP2R5CP, is present at 3p21.3. We also report the genomic organization of both the functional gene and the pseudogene.

Base Sequence↗

Identification of an alternative nucleoside triphosphate: 5'-deoxyadenosylcobinamide phosphate nucleotidyltransferase in Methanobacterium thermoautotrophicum delta H.

Computer analysis of the archaeal genome databases failed to identify orthologues of all of the bacterial cobamide biosynthetic enzymes. Of particular interest was the lack of an orthologue of the bifunctional nucleoside triphosphate (NTP):5'-deoxyadenosylcobinamide kinase/GTP:adenosylcobinamide-phosphate guanylyltransferase enzyme (CobU in Salmonella enterica). This paper reports the identification of an archaeal gene encoding a new nucleotidyltransferase, which is proposed to be the nonorthologous replacement of the S. enterica cobU gene. The gene encoding this nucleotidyltransferase was identified using comparative genome analysis of the sequenced archaeal genomes. Orthologues of the gene encoding this activity are limited at present to members of the domain Archaea. The corresponding ORF open reading frame from Methanobacterium thermoautotrophicum Delta H (MTH1152; referred to as cobY) was amplified and cloned, and the CobY protein was expressed and purified from Escherichia coli as a hexahistidine-tagged fusion protein. This enzyme had GTP:adenosylcobinamide-phosphate guanylyltransferase activity but did not have the NTP:AdoCbi kinase activity associated with the CobU enzyme of S. enterica. NTP:adenosylcobinamide kinase activity was not detected in M. thermoautotrophicum Delta H cell extract, suggesting that this organism may not have this activity. The cobY gene complemented a cobU mutant of S. enterica grown under anaerobic conditions where growth of the cell depended on de novo adenosylcobalamin biosynthesis. cobY, however, failed to restore adenosylcobalamin biosynthesis in cobU mutants grown under aerobic conditions where de novo synthesis of this coenzyme was blocked, and growth of the cell depended on the assimilation of exogenous cobinamide. These data strongly support the proposal that the relevant cobinamide intermediates during de novo adenosylcobalamin biosynthesis are adenosylcobinamide-phosphate and adenosylcobinamide-GDP, not adenosylcobinamide. Therefore, NTP:adenosylcobinamide kinase activity is not required for de novo cobamide biosynthesis.

Amino Acid Sequence↗

Candidate genes in breast cancer revealed by microarray-based comparative genomic hybridization of archived tissue.

Genomic imbalances in 31 formalin-fixed and paraffin-embedded primary tumors of advanced breast cancer were analyzed by microarray-based comparative genomic hybridization (matrix-CGH). A DNA chip was designed comprising 422 mapped genomic sequences including 47 proto-oncogenes, 15 tumor suppressor genes, as well as frequently imbalanced chromosomal regions. Analysis of the data was challenging due to the impaired quality of DNA prepared from paraffin-embedded samples. Nevertheless, using a method for the statistical evaluation of the balanced state for each individual experiment, we were able to reveal imbalances with high significance, which were in good concordance with previous data collected by chromosomal CGH from the same patients. Owing to the improved resolution of matrix-CGH, genomic imbalances could be narrowed down to the level of individual bacterial artificial chromosome and P1-derived artificial chromosome clones. On average 37 gains and 13 losses per tumor cell genome were scored. Gains in more than 30% of the cases were found on 1p, 1q, 6p, 7p, 8q, 9q, 11q, 12q, 17p, 17q, 20q, and 22q, and losses on 6q, 9p, 11q, and 17p. Of the 51 chromosomal regions found amplified by matrix-CGH, only 12 had been identified by chromosomal CGH. Within these 51 amplicons, genome database information defined 112 candidate genes, 44 of which were validated by either PCR amplification of sequence tag sites or DNA sequence analysis.

Breast Neoplasms↗

PCR-analyzed microsatellites of the mouse genome--additional polymorphisms among ten inbred mouse strains.

Eighty sequences from the mouse genome database containing microsatellites (simple sequence repeats) have been analyzed for size variation among ten different inbred strains of mice; 62/80 (77.5%) showed polymorphism of at least three alleles. We have been able to detect all the polymorphisms by agarose gel electrophoresis, often running the gels for up to 3 h. Between individual pairs of mouse strains to be used in chromosomal mapping studies in our laboratory, 35-60% polymorphism occurred. There are potentially enough microsatellites within the mouse and human genome to have a marker at every 1-cM distance. This simple approach will, therefore, continue to be useful in genome mapping studies, leading eventually to high-resolution maps of both the mouse and human genomes; this should allow for physical mapping and cloning of specific genes.

Animals↗

A segment of the apospory-specific genomic region is highly microsyntenic not only between the apomicts Pennisetum squamulatum and buffelgrass, but also with a rice chromosome 11 centromeric-proximal genomic region.

Bacterial artificial chromosome (BAC) clones from apomicts Pennisetum squamulatum and buffelgrass (Cenchrus ciliaris), isolated with the apospory-specific genomic region (ASGR) marker ugt197, were assembled into contigs that were extended by chromosome walking. Gene-like sequences from contigs were identified by shotgun sequencing and BLAST searches, and used to isolate orthologous rice contigs. Additional gene-like sequences in the apomicts' contigs were identified by bioinformatics using fully sequenced BACs from orthologous rice contigs as templates, as well as by interspecies, whole-contig cross-hybridizations. Hierarchical contig orthology was rapidly assessed by constructing detailed long-range contig molecular maps showing the distribution of gene-like sequences and markers, and searching for microsyntenic patterns of sequence identity and spatial distribution within and across species contigs. We found microsynteny between P. squamulatum and buffelgrass contigs. Importantly, this approach also enabled us to isolate from within the rice (Oryza sativa) genome contig Rice A, which shows the highest microsynteny and is most orthologous to the ugt197-containing C1C buffelgrass contig. Contig Rice A belongs to the rice genome database contig 77 (according to the current September 12, 2003, rice fingerprint contig build) that maps proximal to the chromosome 11 centromere, a feature that interestingly correlates with the mapping of ASGR-linked BACs proximal to the centromere or centromere-like sequences. Thus, relatedness between these two orthologous contigs is supported both by their molecular microstructure and by their centromeric-proximal location. Our discoveries promote the use of a microsynteny-based positional-cloning approach using the rice genome as a template to aid in constructing the ASGR toward the isolation of genes underlying apospory.

Cenchrus↗

Converging on a general model of protein evolution.

The availability of high-throughput genomic databases that establish protein dispensability, expression and interaction networks enables rigorous tests of competing models of protein evolution. Recent research utilizing these new data sets shows that protein evolution is more complex than was previously thought. Several variables, including protein dispensability, expression, functional density, and genetic modularity, appear to have independent effects on the evolutionary rate of proteins, suggesting that proteomes have evolved via an assembly of selectional regimes. These results indicate that a general model of protein evolution will emerge as more functional genomic data from a diversity of organisms accumulate.

Evolution, Molecular↗

Towards construction of a high resolution map of the mouse genome using PCR-analysed microsatellites.

Fifty sequences from the mouse genome database containing simple sequence repeats or microsatellites have been analysed for size variation using the polymerase chain reaction and gel electrophoresis. 88% of the sequences, most of which contain the dinucleotide repeat, CA/GT, showed size variations between different inbred strains of mice and the wild mouse, Mus spretus. 62% of sequences had 3 or more alleles. GA/CT and AT/TA-containing sequences were also variable. About half of these size variants were detectable by agarose gel electrophoresis. This simple approach is extremely useful in linkage and genome mapping studies and will facilitate construction of high resolution maps of both the mouse and human genomes.

Animals↗

Proteomic characterization of wheat amyloplasts using identification of proteins by tandem mass spectrometry.

We describe the initial characterization of the wheat amyloplast proteome, consisting of the identification and classification of 171 proteins. Whole amyloplasts and purified amyloplast membranes were prepared from wheat (Triticum aestivum). Protein extracts were examined by one-dimensional and two-dimensional electrophoresis, followed by high performance liquid chromatography-tandem mass spectrometry of separated proteins. Tandem mass spectrometry data of individual peptides was then searched by SEQUEST, using a database containing known protein sequences from both wheat and other homologous cereal crops. Using this approach we identified 108 proteins from whole amyloplasts and 63 proteins from purified amyloplast membranes. The majority of protein identifications were derived from protein sequences from cereal crops other than wheat, for which relatively little gene sequence data is available. The highest percentage of protein identifications obtained from any individual species was 46% of the total number of proteins identified, using sequence data found in our proprietary rice (Oryza sativa) genome database.

Chromatography, High Pressure Liquid↗

NEP1 orthologs encoding necrosis and ethylene inducing proteins exist as a multigene family in Phytophthora megakarya, causal agent of black pod disease on cacao.

Phvytophthora megakarya is a devastating oomycete pathogen that causes black pod disease in cacao. Phytophthora species produce a protein that has a similar sequence to the necrosis and ethylene inducing protein (Nep1) of Fusarium oxysporum. Multiple copies of NEP1 orthologs (PmegNEP) have been identified in P. megakarya and four other Phytophthora species (P. citrophthora, P. capsici, P. palmivora, and P. sojae). Genome database searches confirmed the existence of multiple copies of NEP1 orthologs in P. sojae and P. ramorum. In this study, nine different PmegNEP orthologs from P. megakarya strain Mk-1 were identified and analyzed. Of these nine orthologs, six were expressed in mycelium and in P. megakarya zoospore-infected cacao leaf tissue. The remaining two clones are either regulated differently, or are nonfunctional genes. Sequence analysis revealed that six PmegNEP orthologs were organized in two clusters of three orthologs each in the P. megakarya genome. Evidence is presented for the instability in the P. megakarya genome resulting from duplications, inversions, and fused genes resulting in multiple NEP1 orthologs. Traits characteristic of the Phytophthora genome, such as the clustering of NEP1 orthologs, the lack of CATT and TATA boxes, the lack of introns, and the short distance between ORFs were also observed.

Amino Acid Sequence↗

Life-cycle differentiation in Trypanosoma brucei: molecules and mutants.

Differentiation between bloodstream and tsetse midgut procyclic forms during the life cycle of the African trypanosome is an attractive model for the analysis of stage-regulated events. In particular, this transformation occurs synchronously, there are well-defined markers for stage-regulated processes and cell lines with specific defects in differentiation have been identified. This combination of tools, combined with the developing Trypanosoma brucei genome database is allowing its underlying controls to be investigated at the molecular and cytological levels. This paper examines some recent discoveries that illuminate some of the key events during trypanosome life-cycle progression.

Animals↗

FastR: fast database search tool for non-coding RNA.

The discovery of novel non-coding RNAs has been among the most exciting recent developments in Biology. Yet, many more remain undiscovered. It has been hypothesized that there is in fact an abundance of functional non-coding RNA (ncRNA) with various catalytic and regulatory functions. Computational methods tailored specifically for ncRNA are being actively developed. As the inherent signal for ncRNA is weaker than that for protein coding genes, comparative methods offer the most promising approach, and are the subject of our research. We consider the following problem: Given an RNA sequence with a known secondary structure, efficiently compute all structural homologs (computed as a function of sequence and structural similarity) in a genomic database. Our approach, based on structural filters that eliminate a large portion of the database, while retaining the true homologs allows us to search a typical bacterial database in minutes on a standard PC, with high sensitivity and specificity. This is two orders of magnitude better than current available software for the problem.

Algorithms↗

Transcriptional analysis of the orphan nuclear receptor constitutive androstane receptor (NR1I3) gene promoter: identification of a distal glucocorticoid response element.

The constitutive androstane receptor (CAR, NR1I3) transcriptionally activates cytochrome P450 2B6, 2C9, and 3A4 when activated by xenobiotics, such as phenobarbital. Information on the human CAR promoter was obtained by searching the NCBI human genome database. A contig (NT026945) corresponding to a fragment of chromosome 1q21 was found to contain the complete CAR gene. These data were confirmed using chromosomal in situ hybridization. Both primer extension and 5'-rapid amplification of the cDNA end PCR analysis were carried out to determine the transcriptional start site of human CAR, which was found to be 32 nucleotides downstream of a potential TATA box (CATAAAA). In addition, we found that the 5'-untranslated region of CAR mRNA is 110 nucleotides shorter than previously reported. Using genomic PCR, we amplified and cloned approximately 4.9 kb (-4711/+144) of the CAR gene promoter. The activity of this promoter was measured by transient transfection. Deletion analysis suggested the presence of a glucocorticoid responsive element in its distal region (-4477/-4410). From cotransfection experiments, mutagenesis, and gel shift assays, we identified a glucocorticoid response element at -4447/-4432 that was recognized and transactivated by the human glucocorticoid receptor. Finally, using the chromatin immunoprecipitation assay, we demonstrated that the glucocorticoid receptor binds to the distal region of CAR promoter in cultured hepatocytes only in the presence of dexamethasone. Identification of this functional element provides a rational mechanistic basis for CAR induction by glucocorticoids. CAR appears to be a primary glucocorticoid receptor-response gene.

Cells, Cultured↗

Microbial genomes and "missing" enzymes: redefining biochemical pathways.

A biochemical pathway is the representation of a defined set of substrates, enzyme reactions and products linked together to generate an outcome beneficial to a living cell. Microbial genome sequence data are unparalleled resources for understanding cellular metabolism without the prior definitions imposed by classical biochemistry. Simple analysis of three well-studied biochemical pathways (the tricarboxylic acid cycle, pentose phosphate pathway and glycolysis) from the 17 publicly available microbial genomes has shown that these pathways may rarely occur as previously defined. Therefore, following whole-genome sequencing it has become necessary to redefine the "classical" biochemical steps leading from substrate to end-product for each pathway. Often, unique or alternative reactions appear to be required in order to maintain pathway functionality where expected enzyme reactions (as defined by the presence or absence of the corresponding genes) are "missing". Conversely, such enzymes may be accounted for by: (1) the presence of low sequence similarity or novel genes encoding enzymes performing the same or similar functions, (2) the presence of multienzyme proteins, (3) incorrectly assigned gene identities in genome databases, and (4) known enzyme functions that have yet to be correlated with a gene sequence. Most importantly, the presence of a gene sequence does not necessarily ensure that its corresponding enzyme is actually functional. This may be due to the presence of nonactive remnant genes, evolutionary pressures leading to loss of function, inactivating mutations, the lack of transcription/translation, and post-translational processing. Modifications at the gene and/or functional levels, as well as the possible use of alternative enzymes, must be considered when reconstructing biochemical pathways for fully sequenced microbial genomes.

Bacteria↗

Computational identification, cloning, and characterization of IL-1R9, a novel interleukin-1 receptor-like gene encoded over an unusually large interval of human chromosome Xq22.2-q22.3.

The Interleukin-1 receptor (IL-1R) and Toll signaling pathways share the evolutionarily conserved Toll homology domain (THD), which is a critical component in the signaling cascade of the host defense responses to infection and inflammation. Our initial genomic database searches uncovered a novel THD signature sequence between DNA markers DXS87 and DXS366. The feasibility of subsequently applying a coordinated computational approach, including various exon-finding programs, homology-based searches, and receptor profile searches, in revealing the exons encoding this novel IL-1R family member is described. IL-1R9 shows restricted expression in fetal brain and is highly homologous to IL1RAPL (A. Carrie et al., 1999 Nat. Genet. 23: 25-31), which is reportedly involved in nonsyndromic X-linked mental retardation. These genes are scattered over separate genomic intervals in excess of 1.0 Mb and encode receptors with extended C-terminal tails. In our functional NF-kappaB reporter assays, IL1RAPL, IL-1R9, or versions lacking the extended C-terminal sequences failed in responding either to IL-1 directly or to IL-18 when various permutations of IL-18R ectodomain chimeras were fused to their cytoplasmic domains. Evolutionary sequence analyses reinforce our conclusion that these novel orphan receptors probably form a functionally distinct subset of the IL-1R superfamily.

Amino Acid Motifs↗

Characterization of a Helicobacter pylori vaccine candidate by proteome techniques.

In a previous two-dimensional (2D) gel electrophoretic study of protein antigens of the gastric pathogen, Helicobacter pylori recognized by human sera, one of the highly and consistently reactive antigens, a protein with Mr of approximately 30,000 (Spot 15) seemed to be of special interest because of low yields on N-terminal protein sequencing. This suggested possible N-terminal modification, as the N-terminal sequence analysis of this 30,000 protein (Spot 15) did not provide a definitive match within the H. pylori genomic database. This protein was isolated by 2D polyacrylamide gel electrophoresis, evaluated by liquid chromatography-mass spectrometry, and found to consist of two related species of approximately 28,100 and 26,500. In parallel, the proteins within this spot were digested in situ with the endoprotease Lys-C. Analysis of the Lys-C digest by matrix-assisted laser desorption time-of-flight mass spectrometry, peptide mapping, and sequence analysis was conducted. Comparison of the mass and sequence of the Lys-C peptides with those derived from a H. pylori genomic library identified an open reading frame of approximately 300 base pairs as the source of the Spot 15 protein. This corresponded to HP0175 in the recently reported H. pylori genome sequence, an open reading frame with some homology to Campylobacter jejeuni cell binding protein 2. Mass spectral and sequence analysis indicated that Spot 15 was a processed product generated by proteolytic cleavage at both the carboxy and amino termini of the 34 open reading frame precursor.

Amino Acid Sequence↗

The phylogenetic relationship of the glutamate and pheromone G-protein-coupled receptors in different vertebrate species.

Searches in genomic databases for human, mouse, zebrafish, and pufferfish genes resulted in the identification of more than 180 protein predictions belonging to the glutamate family of G-protein-coupled receptors (GPCRs). Comparison of data sets from the different species showed that most of the receptor subgroups that form the glutamate family are present in both mammalian and bony fish lineage. This finding indicates that these groups share a phylogenetically ancient origin. The present study also shows that the pheromone-receptor subgroup has undergone independent expansions in three of the four species, leaving the human genome completely deprived of all pheromone receptors.

Animals↗

Prediction of many new exons and introns in Plasmodium falciparum chromosome 2.

The current prediction of genes in the Plasmodium falciparum genome database relies upon a limited number of specially developed computer algorithms. We have re-annotated the sequence of chromosome 2 of P. falciparum by a computer-assisted manual analysis, which is described here. Of 161 newly predicted introns, we have experimentally confirmed 98. We regard 110 introns from the previously published analyses as probable, we delete 3, change 26 and add 135. We recognise 214 genes in chromosome 2. We have predicted introns in 121 genes. The increased complexity of gene structure on chromosome 2 is likely to be mirrored by the entire genome.

Algorithms↗