Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Search for potential vaccine candidate open reading frames in the Bacillus anthracis virulence plasmid pXO1: in silico and in vitro screening.

A genomic analysis of the Bacillus anthracis virulence plasmid pXO1, aimed at identifying potential vaccine candidates and virulence-related genes, was carried out. The 143 previously defined open reading frames (ORFs) (R. T. Okinaka, K. Cloud, O. Hampton, A. R. Hoffmaster, K. K. Hill, P. Keim, T. M. Koehler, G. Lamke, S. Kumano, J. Mahillon, D. Manter, Y. Martinez, D. Ricke, R. Svensson, and P. J. Jackson, J. Bacteriol. 181:6509-6515, 1999) were subjected to extensive sequence similarity searches (with the nonredundant and unfinished microbial genome databases), as well as motif, cellular location, and domain analyses. A comparative genomics analysis was conducted with the related genomes of Bacillus subtilis, Bacillus halodurans, and Bacillus cereus and the pBtoxis plasmid of Bacillus thuringiensis var. israeliensis. As a result, the percentage of ORFs with clues about their functions increased from approximately 30% (as previously reported) to more than 60%. The bioinformatics analysis permitted identification of novel genes with putative relevance for pathogenesis and virulence. Based on our analyses, 11 putative proteins were chosen as targets for functional genomics studies. A rapid and efficient functional screening method was developed, in which PCR-amplified full-length linear DNA products of the selected ORFs were transcribed and directly translated in vitro and their immunogenicities were assessed on the basis of their reactivities with hyperimmune anti-B. anthracis antisera. Of the 11 ORFs selected for analysis, 9 were successfully expressed as full-length polypeptides, and 3 of these were found to be antigenic and to have immunogenic potential. The latter ORFs are currently being evaluated to determine their vaccine potential.

Animals↗

Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts.

BACKGROUND: Text-mining can assist biomedical researchers in reducing information overload by extracting useful knowledge from large collections of text. We developed a novel text-mining method based on analyzing the network structure created by symbol co-occurrences as a way to extend the capabilities of knowledge extraction. The method was applied to the task of automatic gene and protein name synonym extraction. RESULTS: Performance was measured on a test set consisting of about 50,000 abstracts from one year of MEDLINE. Synonyms retrieved from curated genomics databases were used as a gold standard. The system obtained a maximum F-score of 22.21% (23.18% precision and 21.36% recall), with high efficiency in the use of seed pairs. CONCLUSION: The method performs comparably with other studied methods, does not rely on sophisticated named-entity recognition, and requires little initial seed knowledge.

Algorithms↗

Genome-derived vaccines.

Vaccine research entered a new era when the complete genome of a pathogenic bacterium was published in 1995. Since then, more than 97 bacterial pathogens have been sequenced and at least 110 additional projects are now in progress. Genome sequencing has also dramatically accelerated: high-throughput facilities can draft the sequence of an entire microbe (two to four megabases) in 1 to 2 days. Vaccine developers are using microarrays, immunoinformatics, proteomics and high-throughput immunology assays to reduce the truly unmanageable volume of information available in genome databases to a manageable size. Vaccines composed by novel antigens discovered from genome mining are already in clinical trials. Within 5 years we can expect to see a novel class of vaccines composed by genome-predicted, assembled and engineered T- and Bcell epitopes. This article addresses the convergence of three forces--microbial genome sequencing, computational immunology and new vaccine technologies--that are shifting genome mining for vaccines onto the forefront of immunology research.

Animals↗

Allele frequency of D1S191 microsatellite locus in Japanese people.

The allele frequency of D1S191 microsatellite locus was analyzed in 398 normal tissues of Japanese adults. The frequency of D1S191 microsatellite polymorphism was 90.2% (359/398). The size of D1S191 PCR fragments ranged from 147 bp to 169 bp. The most frequent allele in the Japanese subjects was 161 bp (34.1%) followed by 159 bp (27.9%), 163 bp (17.7%), and 157 bp (12.4%) fragments. The observed allelic distribution of this microsatellite in the Japanese subjects was similar to that of European Caucasians deposited in the Genome Database.

Adult↗

The human genome, implications for oral health and diseases, and dental education.

We are living in an extraordinary time in human history punctuated by the convergence of major scientific and technological progress in the physical, chemical, and biological ways of knowing. Equally extraordinary are the sparkling intellectual developments at the interface between fields of study. One major example of an emerging influence on the future of oral health education is at the interface between the human genome, information technology, and biotechnology with miniaturizations (nanotechnology), suggesting new oral health professional competencies for a new century. A great deal has recently been learned from human and non-human genomics. Genome databases are being "mined" to prompt hypothesis-driven "postgenomic" or functional genomic science in microbial models such as Candida albicans related to oral candidiasis and in human genomics related to biological processes found in craniofacial, oral, and dental diseases and disorders. This growing body of knowledge is already providing the gene content of many oral microbial and human genomes and the knowledge of genetic variants or polymorphisms related to disease, disease progression, and disease response to therapeutics (pharmacogenomics). The knowledge base from human and non-human genomics, functional genomics, biotechnology, and associated information technologies is serving to revolutionize oral health promotion, risk assessment using biomarkers and disease prevention, diagnostics, treatments, and the full range of therapeutics for craniofacial, oral, and dental diseases and disorders. Education, training, and research opportunities are already transforming the curriculum and pedagogy for undergraduate science majors, predoctoral health professional programs, residency and specialty programs, and graduate programs within the health professions. In the words of Bob Dylan, "the times they are a-changing."

Biotechnology↗

A novel genomics approach for the identification of drug targets in pathogens, with special reference to Pseudomonas aeruginosa.

Complete genome sequences of several pathogenic bacteria have been determined, and many more such projects are currently under way. While these data potentially contain all the determinants of host-pathogen interactions and possible drug targets, computational tools for selecting suitable candidates for further experimental analyses are currently limited. Detection of bacterial genes that are non-homologous to human genes, and are essential for the survival of the pathogen represents a promising means of identifying novel drug targets. We have used three-way genome comparisons to identify essential genes from Pseudomonas aeruginosa. Our approach identified 306 essential genes that may be considered as potential drug targets. The resultant analyses are in good agreement with the results of systematic gene deletion experiments. This approach enables rapid potential drug target identification, thereby greatly facilitating the search for new antibiotics. These results underscore the utility of large genomic databases for in silico systematic drug target identification in the post-genomic era.

Anti-Bacterial Agents↗

GARSA: genomic analysis resources for sequence annotation.

SUMMARY: Growth of genome data and analysis possibilities have brought new levels of difficulty for scientists to understand, integrate and deal with all this ever-increasing information. In this scenario, GARSA has been conceived aiming to facilitate the tasks of integrating, analyzing and presenting genomic information from several bioinformatics tools and genomic databases, in a flexible way. GARSA is a user-friendly web-based system designed to analyze genomic data in the context of a pipeline. EST and GGS data can be analyzed using the system since it accepts (1) chromatograms, (2) download of sequences from GenBank, (3) Fasta files stored locally or (4) a combination of all three. Quality evaluation of chromatograms, vector removing and clusterization are easily performed as part of the pipeline. A number of local and customizable Blast and CDD analyses can be performed as well as Interpro, complemented with phylogeny analyses. GARSA is being used for the analyses of Trypanosoma vivax (GSS and EST), Trypanosoma rangeli (GSS, EST and ORESTES), Bothrops jararaca (EST), Piaractus mesopotamicus (EST) and Lutzomyia longipalpis (EST). AVAILABILITY: The GARSA system is freely available under GPL license (http://www.biowebdb.org/garsa/). For download requests visit http://www.biowebdb.org/garsa/ or contact Dr Alberto Dávila.

Animals↗

Prediction of a novel RNA 2'-O-ribose methyltransferase subfamily encoded by the Escherichia coli YgdE open reading frame and its orthologs.

The amino acid sequence of the RNA 2'-O-ribose methyltranserase RrmJ was used as a probe for detecting putative homologs through iterative searches of genomic databases. We found a previously unannotated YgdE open reading frame (ORF) in the genome sequences of Escherichia coli and other gamma-Proteobacteria, which shares key features with RrmJ, despite the mutual sequence similarity of these proteins is relatively low. The predicted structural compatibility and the conservation of all functionally important residues between RrmJ and YgdE strongly suggests that the newly identified methyltranserase also modifies 2'-OH groups of ribose. The N-terminal region of YgdE, which has no counterpart in RrmJ, is predicted to form an independent domain, possibly involved in target recognition.

Amino Acid Sequence↗

GenoMycDB: a database for comparative analysis of mycobacterial genes and genomes.

Several databases and computational tools have been created with the aim of organizing, integrating and analyzing the wealth of information generated by large-scale sequencing projects of mycobacterial genomes and those of other organisms. However, with very few exceptions, these databases and tools do not allow for massive and/or dynamic comparison of these data. GenoMycDB (http://www.dbbm.fiocruz.br/GenoMycDB) is a relational database built for large-scale comparative analyses of completely sequenced mycobacterial genomes, based on their predicted protein content. Its central structure is composed of the results obtained after pair-wise sequence alignments among all the predicted proteins coded by the genomes of six mycobacteria: Mycobacterium tuberculosis (strains H37Rv and CDC1551), M. bovis AF2122/97, M. avium subsp. paratuberculosis K10, M. leprae TN, and M. smegmatis MC2 155. The database stores the computed similarity parameters of every aligned pair, providing for each protein sequence the predicted subcellular localization, the assigned cluster of orthologous groups, the features of the corresponding gene, and links to several important databases. Tables containing pairs or groups of potential homologs between selected species/strains can be produced dynamically by user-defined criteria, based on one or multiple sequence similarity parameters. In addition, searches can be restricted according to the predicted subcellular localization of the protein, the DNA strand of the corresponding gene and/or the description of the protein. Massive data search and/or retrieval are available, and different ways of exporting the result are offered. GenoMycDB provides an on-line resource for the functional classification of mycobacterial proteins as well as for the analysis of genome structure, organization, and evolution.

Bacterial Proteins↗

AluGene: a database of Alu elements incorporated within protein-coding genes.

Alu elements are short interspersed elements (SINEs) approximately 300 nucleotides in length. More than 1 million Alus are found in the human genome. Despite their being genetically functionless, recent findings suggest that Alu elements may have a broad evolutionary impact by affecting gene structures, protein sequences, splicing motifs and expression patterns. Because of these effects, compiling a genomic database of Alu sequences that reside within protein-coding genes seemed a useful enterprise. Presently, such data are limited since the structural and positional information on genes and Alu sequences are scattered throughout incompatible and unconnected databases. AluGene (http://Alugene.tau.ac.il/) provides easy access to a complete Alu map of the human genome, as well as Alu-associated information. The Alu elements are annotated with respect to coding region and exon/intron location. This design facilitates queries on Alu sequences, locations, as well as motifs and compositional properties via a one-stop search page.

Alu Elements↗

In silico characterization of the family of PARP-like poly(ADP-ribosyl)transferases (pARTs).

BACKGROUND: ADP-ribosylation is an enzyme-catalyzed posttranslational protein modification in which mono(ADP-ribosyl)transferases (mARTs) and poly(ADP-ribosyl)transferases (pARTs) transfer the ADP-ribose moiety from NAD onto specific amino acid side chains and/or ADP-ribose units on target proteins. RESULTS: Using a combination of database search tools we identified the genes encoding recognizable pART domains in the public genome databases. In humans, the pART family encompasses 17 members. For 16 of these genes, an orthologue exists also in the mouse, rat, and pufferfish. Based on the degree of amino acid sequence similarity in the catalytic domain, conserved intron positions, and fused protein domains, pARTs can be divided into five major subgroups. All six members of groups 1 and 2 contain the H-Y-E trias of amino acid residues found also in the active sites of Diphtheria toxin and Pseudomonas exotoxin A, while the eleven members of groups 3 - 5 carry variations of this motif. The pART catalytic domain is found associated in Lego-like fashion with a variety of domains, including nucleic acid-binding, protein-protein interaction, and ubiquitylation domains. Some of these domain associations appear to be very ancient since they are observed also in insects, fungi, amoebae, and plants. The recently completed genome of the pufferfish T. nigroviridis contains recognizable orthologues for all pARTs except for pART7. The nearly completed albeit still fragmentary chicken genome contains recognizable orthologues for twelve pARTs. Simpler eucaryotes generally contain fewer pARTs: two in the fly D. melanogaster, three each in the mosquito A. gambiae, the nematode C. elegans, and the ascomycete microfungus G. zeae, six in the amoeba E. histolytica, nine in the slime mold D. discoideum, and ten in the cress plant A. thaliana. GenBank contains two pART homologues from the large double stranded DNA viruses Chilo iridescent virus and Bacteriophage Aeh1 and only a single entry (from V. cholerae) showing recognizable homology to the pART-like catalytic domains of Diphtheria toxin and Pseudomonas exotoxin A. CONCLUSION: The pART family, which encompasses 17 members in the human and 16 members in the mouse, can be divided into five subgroups on the basis of sequence similarity, phylogeny, conserved intron positions, and patterns of genetically fused protein domains.

Adenosine Diphosphate↗

Synechocystis sp. PCC 6803 - a useful tool in the study of the genetics of cyanobacteria.

The cyanobacterium Synechocystis sp. PCC 6803 was the first phototrophic organism to be fully sequenced. The genomic sequence has revealed the structure of the genome and its gene constituents (3167 genes), as well as the relative map positions of each gene. The functions of nearly half of the genes has been deduced using similarity searches. The genome sequence has also allowed for the implementation of systematic strategies to study gene function and the mechanisms of gene regulation on a genome-wide level. Two genome databases, CyanoBase and CyanoMutants, have been established and act as a central repository for information on gene structure and gene function, respectively. As a result of the genome sequencing and the establishment of these databases, Synechocystis sp. PCC 6803 provides an extremely versatile and easy model to study the genetic systems of photosynthetic organisms.

Journal Article↗

The yeast Saccharomyces cerevisiae YDL112w ORF encodes the putative 2'-O-ribose methyltransferase catalyzing the formation of Gm18 in tRNAs.

The protein sequences of three known RNA 2'-O-ribose methylases were used as probes for detecting putative homologs through iterative searches of genomic databases. We have identified 45 new positive Open Reading Frames (ORFs), mostly in prokaryotic genomes. Five complete eukaryotic ORFs were also detected, among which was a single ORF (YDL112w) in the yeast Saccharomyces cerevisiae genome. After genetic depletion of YDL112w, we observed a specific defect in tRNA ribose methylation, with the complete disappearance of Gm18 in all tRNAs that naturally contain this modification, whereas other tRNA ribose methylations and the complex pattern of rRNA ribose methylations were not affected. The tRNA G18 methylation defect was suppressed by transformation of the disrupted strain with a plasmid allowing expression of YDL112wp. The formation of Gm18 on an in vitro transcript of a yeast tRNASer naturally containing this methylation, which was efficiently catalyzed by cell-free extracts from the wild-type yeast strain, did not occur with extracts from the disrupted strain. The protein encoded by the YDL112w ORF, termed Trm3 (tRNA methylation), is therefore likely to be the tRNA (Gm18) ribose methylase. In in vitro assays, its activity is strongly dependent on tRNA architecture. Trm3p, the first putative tRNA ribose methylase identified in an eukaryotic organism, is considerably larger than its Escherichia coli functional homolog spoU (1,436 amino acids vs. 229 amino acids), or any known or putative prokaryotic RNA ribose methyltransferase. Homologs found in human (TRP-185 protein), Caenorhabditis elegans and Arabidopsis thaliana also exhibit a very long N-terminal extension not related to any protein sequence in databases.

Amino Acid Sequence↗

Cloning and expression of the S-adenosylmethionine decarboxylase gene of Neurospora crassa and processing of its product.

S-adenosylmethionine decarboxylase (AdoMetDC) catalyzes the formation of decarboxylated AdoMetDC, a precursor of the polyamines spermidine and spermine. The enzyme is derived from a proenzyme by autocatalytic cleavage. We report the cloning and regulation of the gene for AdoMetDC in Neurospora crassa, spe-2, and the effect of putrescine on enzyme maturation and activity. The gene was cloned from a genomic library by complementation of a spe-2 mutant. Like other AdoMetDCs, that of Neurospora is derived by cleavage of a proenzyme. The deduced sequence of the Neurospora proenzyme (503 codons) is over 100 codons longer than any other AdoMetDC sequence available in genomic databases. The additional amino acids are found only in the AdoMetDC of another fungus, Aspergillus nidulans, a cDNA for which we also sequenced. Despite the conserved processing site and four acidic residues required for putrescine stimulation of human proenzyme processing, putrescine has no effect on the rate (t0.5 approximately 10 min) of processing of the Neurospora gene product. However, putrescine is absolutely required for activity of the Neurospora enzyme (K0.5 approximately 100 microM). The abundance of spe-2 mRNA and enzyme activity is regulated 2- to 4-fold by spermidine.

Adenosylmethionine Decarboxylase↗

Identification of paralogous HERV-K LTRs on human chromosomes 3, 4, 7 and 11 in regions containing clusters of olfactory receptor genes.

A locus harboring a human endogenous retroviral LTR (long terminal repeat) was mapped on the short arm of human chromosome 7 (7p22), and its evolutionary history was investigated. Sequences of two human genome fragments that were homologous to the LTR-flanking sequences were found in human genome databases: (1) an LTR-containing DNA fragment from region 3p13 of the human genome, which includes clusters of olfactory receptor genes and pseudogenes; and (2) a fragment of region 21q22.1 lacking LTR sequences. PCR analysis demonstrated that LTRs with highly homologous flanking sequences could be found in the genomes of human, chimp, gorilla, and orangutan, but were absent from the genomes of gibbon and New World monkeys. A PCR assay with a primer set corresponding to the sequence from human Chr 3 allowed us to detect LTR-containing paralogous sequences on human chromosomes 3, 4, 7, and 11. The divergence times for the LTR-flanking sequences on chromosomes 3 and 7, and the paralogous sequence on chromosome 21, were evaluated and used to reconstruct the order of duplication events and retroviral insertions. (1) An initial duplication event that occurred 14-17 Mya and before LTR insertion - produced two loci, one corresponding to that located on Chr 21, while the second was the ancestor of the loci on chromosomes 3 and 7. (2) Insertion of the LTR (most probably as a provirus) into this ancestral locus took place 13 Mya. (3) Duplication of the LTR-containing ancestral locus occurred 11 Mya, forming the paralogous modern loci on Chr 3 and 7.

Chromosome Mapping↗

Molecular cloning and pharmacological characterization of the guinea pig 5-HT1E receptor.

The human 5-HT(1E) receptor gene was cloned more than a decade ago. Little is known about its function, and there have been no reports of its existence in the genome of small laboratory animals. In this study, attempts to clone the 5-HT(1E) gene from the rat and mouse were unsuccessful. In fact, a search of the mouse genome database revealed that the 5-HT(1E) receptor gene is missing from the mouse genome. However, the 5-HT(1E) gene was cloned from guinea pig genomic DNA and was characterized. The guinea pig 5-HT(1E) receptor gene encodes a protein of 365 amino acids. It shares 88% (nucleic acid) and 95% (amino acid) homology with the human receptor. The guinea pig 5-HT(1E) receptor showed similar pharmacology to the human 5-HT(1E) receptor in radioligand binding assays. Serotonin (5-hydroxytryptamine, 5-HT) dose-dependently stimulated [35S]GTPgammaS binding to the guinea pig 5-HT(1E) receptor with an EC(50) of 13.6+/-1.92 nM, similar to that of the human 5-HT(1E) receptor (13.7+/-1.78 nM). Activation of the guinea pig 5-HT(1E) receptor was also achieved by ergonovine, alpha-methyl-5-HT, 1-naphthylpiperazine, methysergide, tryptamine, and 1-(2,5-dimethoxy-4-iodophenyl)-2-aminopropane (DOI). Methiothepin exhibited antagonist activity. Quantitative real-time polymerase chain reaction (qRT-PCR) analysis showed that 5-HT(1E) mRNA was present in the guinea pig brain with the greatest abundance in the hippocampus, followed by the olfactory bulb. Lower levels were detected in the cortex, thalamus, pons, hypothalamus, midbrain, striatum, and cerebellum. Our current study marks the first identification of the 5-HT(1E) receptor gene in a commonly used laboratory animal species. This finding should allow the elucidation of the receptor's role(s) in the complex coordination of central serotonergic effects.

Amino Acid Sequence↗

Mining nematode genome data for novel drug targets.

Expressed sequence tag projects have currently produced over 400 000 partial gene sequences from more than 30 nematode species and the full genomic sequences of selected nematodes are being determined. In addition, functional analyses in the model nematode Caenorhabditis elegans have addressed the role of almost all genes predicted by the genome sequence. This recent explosion in the amount of available nematode DNA sequences, coupled with new gene function data, provides an unprecedented opportunity to identify pre-validated drug targets through efficient mining of nematode genomic databases. This article describes the various information sources available and strategies that can expedite this process.

Animals↗