Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Pseudomonas aeruginosa and a proteomic approach to bacterial pathogenesis.

Pseudomonas aeruginosa is a Gram-negative bacterium that is ubiquitous in the environment and can cause a variety of diseases in compromised patients. The genome of P. aeruginosa strain PAO1 has been reported to contain 5570 potential proteins. The value of this genomic database is that new proteins can be recognized to use as diagnostic markers, novel drug targets, and to better understand the physiology of this organism. However, similar to what has been observed in other sequenced bacterial genomes, approximately one third of the potential proteins have no known function. This is somewhat surprising given the long-standing interest in P. aeruginosa as an opportunistic pathogen. Obviously new tools, in addition to sequence similarity analysis, are needed to determine the role of these proteins. Proteomics using two-dimensional gel electrophoresis followed by mass spectrometry to detect and identify P. aeruginosa proteins represents a novel approach to address this gap.

Animals↗

The Portable Dictionary of the Mouse Genome: a personal database for gene mapping and molecular biology.

The Portable Dictionary of the Mouse Genome is a database for personal computers that contains information on approximately 10,000 loci in the mouse, along with data on homologs in several other mammalian species, including human, rat, cat, cow, and pig. Key features of the dictionary are its compact size, its network independence, and the ability to convert the entire dictionary to a wide variety of common application programs. Another significant feature is the integration of DNA sequence accession data. Loci in the dictionary can be rapidly resorted by chromosomal position, by type, by human homology, or by gene effect. The dictionary provides an accessible, easily manipulated set of data that has many uses--from a quick review of loci and gene nomenclature to the design of experiments and analysis of results. The Portable Dictionary is available in several formats suitable for conversion to different programs and computer systems. It can be obtained on disk or from Internet Gopher servers (mickey.utmen.edu or anat4.utmen.edu), an anonymous FTP site (nb.utmem.edu in the directory pub/genedict), and a World Wide Web server (http://mickey.utmem.edu/front.html).

Animals↗

Candidate gene database and transcript map for peach, a model species for fruit trees.

Peach (Prunus persica) is a model species for the Rosaceae, which includes a number of economically important fruit tree species. To develop an extensive Prunus expressed sequence tag (EST) database for identifying and cloning the genes important to fruit and tree development, we generated 9,984 high-quality ESTs from a peach cDNA library of developing fruit mesocarp. After assembly and annotation, a putative peach unigene set consisting of 3,842 ESTs was defined. Gene ontology (GO) classification was assigned based on the annotation of the single "best hit" match against the Swiss-Prot database. No significant homology could be found in the GenBank nr databases for 24.3% of the sequences. Using core markers from the general Prunus genetic map, we anchored bacterial artificial chromosome (BAC) clones on the genetic map, thereby providing a framework for the construction of a physical and transcript map. A transcript map was developed by hybridizing 1,236 ESTs from the putative peach unigene set and an additional 68 peach cDNA clones against the peach BAC library. Hybridizing ESTs to genetically anchored BACs immediately localized 11.2% of the ESTs on the genetic map. ESTs showed a clustering of expressed genes in defined regions of the linkage groups. [The data were built into a regularly updated Genome Database for Rosaceae (GDR), available at (http://www.genome.clemson.edu/gdr/).].

Breeding↗

Tnfrsf13c (Baffr) is mis-expressed in tumors with murine leukemia virus insertions at Lvis22.

In susceptible strains of mice, leukemia is caused by the somatic integration of murine leukemia retroviruses into the host genome. Integration sites that are common to several tumors are likely to affect genes that are important in oncogenesis. Here we present the analysis of a common site of retroviral integration on mouse chromosome 15, which includes the genomic structure of three genes near the integration site. One of the genes misexpressed at the insertion site has recently been characterized as a B-cell receptor, Tnfrsf13c (formerly Baffr), indicating that this approach is useful in defining genes that function in lymphocyte development and tumor progression. Current genome databases provide powerful resources for the rapid identification of genes at common proviral insertion sites. The characterization of these genes in tumor samples will allow a function to be assigned to many novel loci identified by the genome sequencing projects.

Amino Acid Sequence↗

The human gene mutation database.

The Human Gene Mutation Database (HGMD) represents a comprehensive core collection of data on published germline mutations in nuclear genes underlying human inherited disease. By September 1997, the database contained nearly 12 000 different lesions in a total of 636 different genes, with new entries currently accumulating at a rate of over 2000 per annum. Although originally established for the scientific study of mutational mechanisms in human genes, HGMD has acquired a much broader utility to researchers, physicians and genetic counsellors so that it was made publicly available at http://uwcm.ac.uk/uwcm/mg/hgmd0.html in April 1996. Mutation data in HGMD are accessible on the basis of every gene being allocated one web page per mutation type, if data of that type are present. Meaningful integration with phenotypic, structural and mapping information has been accomplished through bi-directional links between HGMD and both the Genome Database (GDB) and Online Mendelian Inheritance in Man (OMIM), Baltimore, USA. Hypertext links have also been established to Medline abstracts through Entrez , and to a collection of 458 reference cDNA sequences also used for data checking. Being both comprehensive and fully integrated into the existing bioinformatics structures relevant to human genetics, HGMD has established itself as the central core database of inherited human gene mutations.

Computer Communication Networks↗

Bacterial pathogen genomics and vaccines.

Infectious diseases remain a major cause of deaths and disabilities in the world, the majority of which are caused by bacteria. Although immunisation is the most cost effective and efficient means to control microbial diseases, vaccines are not yet available to prevent many major bacterial infections. Examples include dysentery (shigellosis), gonorrhoea, trachoma, gastric ulcers and cancer (Helicobacter pylori). Improved vaccines are needed to combat some diseases for which current vaccines are inadequate. Tuberculosis, for example, remains rampant throughout most countries in the world and represents a global emergency heightened by the pandemic of HIV. The availability of complete genome sequences has dramatically changed the opportunities for developing novel and improved vaccines and facilitated the efficiency and rapidity of their development. Complete genomic databases provide an inclusive catalogue of all potential candidate vaccines for any bacterial pathogen. In conjunction with adjunct technologies, including bioinformatics, random mutagenesis, microarrays, and proteomics, a systematic and comprehensive approach to identifying vaccine discovery can be undertaken. Genomics must be used in conjunction with population biology to ensure that the vaccine can target all pathogenic strains of a species. A proof in principle of the utility of genomics is provided by the recent exploitation of the complete genome sequence of Neisseria meningitidis group B.

Animals↗

Can SINEs: a family of tRNA-derived retroposons specific to the superfamily Canoidea.

A repetitive element of approximately 200 bp was cloned from harbour seal (Phoca vitulina concolour) genomic DNA. The sequence of the element revealed putative RNA polymerase III control boxes, a poly A tail and direct terminal repeats characteristic of SINEs. Sequence and secondary structural similarities suggest that the SINE is derived from a tRNA, possibly tRNA-alanine. Southern blot analysis indicated that the element is predominately dispersed in unique regions of the seal genome, but may also be present in other repetitive sequences, such as tandemly arrayed satellite DNA. Based on slot-blot hybridization analysis, we estimate that 1.3 x 10(6) copies of the SINE are present in the harbour seal genome; SINE copy number based on the number of clones isolated from a size-selected library, however, is an order of magnitude lower (1-3 x 10(5) copies), an estimate consistent with the abundance of SINEs in other mammalian genomes. Database searches found similar sequences have been isolated from dog (Canis familiaris) and mink (Mustela vison). These, and the seal SINE sequences are characterized by an internal CT dinucleotide microsatellite in the tRNA-unrelated region. Hybridization of genomic DNA from representative species of a wide range of mammalian orders to an oligonucleotide (30mer) probe complementary to a conserved region of the SINE confirmed that the element is unique to carnivores of the superfamily Canoidea.

Animals↗

Development of a Fusarium graminearum Affymetrix GeneChip for profiling fungal gene expression in vitro and in planta.

Recently the genome sequences of several filamentous fungi have become available, providing the opportunity for large-scale functional analysis including genome-wide expression analysis. We report the design and validation of the first Affymetrix GeneChip microarray based on the entire genome of a filamentous fungus, the ascomycetous plant pathogen Fusarium graminearum. To maximize the likelihood of representing all putative genes (approximately 14,000) on the array, two distinct sets of automatically predicted gene calls were used and integrated into the online F. graminearum Genome DataBase. From these gene sets, a subset of calls was manually annotated and a non-redundant extract of all calls together with additional EST sequences and controls were submitted for GeneChip design. Experiments were conducted to test the performance of the F. graminearum GeneChip. Hybridization experiments using genomic DNA demonstrated the usefulness of the array for experimentation with F. graminearum and at least four additional pathogenic Fusarium species. Differential transcript accumulation was detected using the F. graminearum GeneChip with treatments derived from the fungus grown in culture under three nutritional regimes and in comparison with fungal growth in infected barley. The ability to detect fungal genes in planta is surprisingly sensitive even without efforts to enrich for fungal transcripts. The Plant Expression Database (PLEXdb, http://www.plexdb.org) will be used as a public repository for raw and normalized expression data from the F. graminearum GeneChip. The F. graminearum GeneChip will help to accelerate exploration of the pathogen-host pathways that may involve interactions between pathogenicity genes in the fungus and disease response in the plant.

Computational Biology↗

Genomically linked cellular protein databases derived from two-dimensional polyacrylamide gel electrophoresis.

In its most useful form a cellular protein database should be genomically based, because it is the genome which determines both the total number of proteins a cell can make and the particular ones that will be made under any given condition. Such a database should trace each protein back to its structural gene, and should account for every structural gene of a cell. Recent advances in molecular biology greatly facilitate the construction of such gene-protein databases. The mapping of genes of unidentified proteins resolved from total cell extracts on two-dimensional gels can now be accomplished by largely biochemical methods, without the necessity of isolating mutants or performing genetic crosses. Other techniques permit one to search gels for the product of any newly discovered gene (or open reading frame) suspected of encoding a protein. Consequently, gene-protein indices can be built independently and simultaneously from either direction--deducing the genetic map from the protein pattern, or finding the protein pattern from information encoded in the genome. A database of this sort is being constructed for the bacterium, Escherichia coli. Given the current pace of DNA nucleotide sequencing, the development of total gene-protein indices for a variety of cells can be anticipated in the near future.

Amino Acids↗

The SNF2 domain protein family in higher vertebrates displays dynamic expression patterns in Xenopus laevis embryos.

All eukaryotes share a common nuclear infrastructure, in which DNA is packaged into nucleosomal chromatin. Its functional states, in particular the accessibility of the chromatin fiber to trans-acting factors, are determined by two classes of evolutionarily conserved enzymes, i.e. histone modifying enzymes and ATP-driven nucleosome remodeling machines. Browsing the annotated human genome database, we establish here a family of SNF2-like nuclear ATPases, which are the core enzymatic subunits of chromatin remodeling protein complexes. Homologues of those human genes are also to a large extent found in the Xenopus laevis genome, indicating a high degree of sequence conservation of this family among vertebrates. Expression analyses of the ATPase family of proteins reveal stage- and tissue-specific domains of peak RNA expression during early frog embryogenesis. These dynamic expression profiles suggest specific functional requirements for individual members of this family throughout early stages of vertebrate development.

Adenosine Triphosphatases↗

Analysis of cytoskeletal and motility proteins in the sea urchin genome assembly.

The sea urchin embryo is a classical model system for studying the role of the cytoskeleton in such events as fertilization, mitosis, cleavage, cell migration and gastrulation. We have conducted an analysis of gene models derived from the Strongylocentrotus purpuratus genome assembly and have gathered strong evidence for the existence of multiple gene families encoding cytoskeletal proteins and their regulators in sea urchin. While many cytoskeletal genes have been cloned from sea urchin with sequences already existing in public databases, genome analysis reveals a significantly higher degree of diversity within certain gene families. Furthermore, genes are described corresponding to homologs of cytoskeletal proteins not previously documented in sea urchins. To illustrate the varying degree of sequence diversity that exists within cytoskeletal gene families, we conducted an analysis of genes encoding actins, specific actin-binding proteins, myosins, tubulins, kinesins, dyneins, specific microtubule-associated proteins, and intermediate filaments. We conducted ontological analysis of select genes to better understand the relatedness of urchin cytoskeletal genes to those of other deuterostomes. We analyzed developmental expression (EST) data to confirm the existence of select gene models and to understand their differential expression during various stages of early development.

Animals↗

Comparative analysis of methodologies for the detection of horizontally transferred genes: a reassessment of first-order Markov models.

With the advent of larger genome databases detection of horizontal gene transfer events has been transformed into an increasingly important issue. Here we present a simple theoretical analysis based on the in silico artificial addition of known foreign genes from different prokaryotic groups into the genome of Escherichia coli K12 MG1655. Using this dataset as a control, we have tested the efficiency of four methodologies commonly employed to detect HTG (Horizontally transferred genes), which are based on (a) the codon adaptation index, codon usage, and GC percentage (CAI/GC); (b) a distributional profile (DP) approach made by a gene search in the closely related phylogenetic genomes; (c) a Bayesian model (BM); and (d) a first-order Markov model (MM). All methods exhibit limitations although, as shown here, the BM and the MM are better approximations. Moreover, the MM has demonstrated a more accurate rate of detections when genes from closely related organisms are evaluated. The application of the MM to detect recently transferred genes in the genomes of E. coli strains K12 MG1655, O157 EDL933, and Salmonella typhimurium, shows that these organisms have undergone a rather significant amount of HTG, most of which appear to be pseudogenes. Few of these sequences that have undergone HGT appear to have well defined functions and may be involved in the organism's adaptation.

Computer Simulation↗

The year of the worm.

Developmental biology has almost come full circle. Initially aimed at description at the organismal level, in the last 25 years it has zoomed in on individual genes that are involved in specific steps in development. Now, complete genome sequences are becoming available--and to gain a full understanding of the relevance of the complete genome, experimental developmental biology will hold centre stage again, but now armed with large genome databases, and with a new set of refined genetic tools. The first multicellular organism to be sequenced is the nematode C. elegans. This review aims to recognise some new avenues in C. elegans experimental biology that are opened by the genome sequence.

Amino Acid Sequence↗

Identification of urocortin III, an additional member of the corticotropin-releasing factor (CRF) family with high affinity for the CRF2 receptor.

The corticotropin-releasing factor (CRF) family of neuropeptides includes the mammalian peptides CRF, urocortin, and urocortin II, as well as piscine urotensin I and frog sauvagine. The mammalian peptides signal through two G protein-coupled receptor types to modulate endocrine, autonomic, and behavioral responses to stress, as well as a range of peripheral (cardiovascular, gastrointestinal, and immune) activities. The three previously known ligands are differentially distributed anatomically and have distinct specificities for the two major receptor types. Here we describe the characterization of an additional CRF-related peptide, urocortin III, in the human and mouse. In searching the public human genome databases we found a partial expressed sequence tagged (EST) clone with significant sequence identity to mammalian and fish urocortin-related peptides. By using primers based on the human EST sequence, a full-length human clone was isolated from genomic DNA that encodes a protein that includes a predicted putative 38-aa peptide structurally related to other known family members. With a human probe, we then cloned the mouse ortholog from a genomic library. Human and mouse urocortin III share 90% identity in the 38-aa putative mature peptide. In the peptide coding region, both human and mouse urocortin III are 76% identical to pufferfish urocortin-related peptide and more distantly related to urocortin II, CRF, and urocortin from other mammalian species. Mouse urocortin III mRNA expression is found in areas of the brain including the hypothalamus, amygdala, and brainstem, but is not evident in the cerebellum, pituitary, or cerebral cortex; it is also expressed peripherally in small intestine and skin. Urocortin III is selective for type 2 CRF receptors and thus represents another potential endogenous ligand for these receptors.

Amino Acid Sequence↗

TXTGate: profiling gene groups with text-based information.

We implemented a framework called TXTGate that combines literature indices of selected public biological resources in a flexible text-mining system designed towards the analysis of groups of genes. By means of tailored vocabularies, term- as well as gene-centric views are offered on selected textual fields and MEDLINE abstracts used in LocusLink and the Saccharomyces Genome Database. Subclustering and links to external resources allow for in-depth analysis of the resulting term profiles.

Animals↗

'Oming in on schistosomes: prospects and limitations for post-genomics.

The recent release of version 3 of the Schistosoma mansoni genome assembly has made a wealth of information available to researchers. Here, progress made in schistosome genomics and post-genomics is considered. The current status of knowledge about the genome, transcriptome, proteome, glycome and immunome is summarized and recent publications briefly reviewed. The prospects for advances in understanding schistosome biology are highlighted. Most importantly, the limitations (which are mostly technical) that need to be addressed before the full potential of the genome database(s) can be realized are emphasized.

Animals↗

TRIPLES: a database of gene function in Saccharomyces cerevisiae.

Using a novel multipurpose mini-transposon, we have generated a collection of defined mutant alleles for the analysis of disruption phenotypes, protein localization, and gene expression in Saccharomyces cerevisiae. To catalog this unique data set, we have developed TRIPLES, a Web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces. Encompassing over 250 000 data points, TRIPLES provides convenient access to information from nearly 7800 transposon-mutagenized yeast strains; within TRIPLES, complete data reports of each strain may be viewed in table format, or if desired, downloaded as tab-delimited text files. Each report contains external links to corresponding entries within the Saccharomyces Genome Database and International Nucleic Acid Sequence Data Library (GenBank). Unlike other yeast databases, TRIPLES also provides on-line order forms linked to each clone report; users may immediately request any desired strain free-of-charge by submitting a completed form. In addition to presenting a wealth of information for over 2300 open reading frames, TRIPLES constitutes an important medium for the distribution of useful reagents throughout the yeast scientific community. Maintained by the Yale Genome Analysis Center, TRIPLES may be accessed at http://ycmi.med.yale.edu/ygac/triples.htm

DNA Transposable Elements↗

DNA Data Bank of Japan (DDBJ) in XML.

The DNA Data Bank of Japan (DDBJ, http://www.ddbj.nig.ac.jp) has collected and released more entries and bases than last year. This is mainly due to large-scale submissions from Japanese sequencing teams on mouse, rice, chimpanzee, nematoda and other organisms. The contributions of DDBJ over the past year are 17.3% (entries) and 10.3% (bases) of the combined outputs of the International Nucleotide Sequence Databases (INSD). Our complete genome sequence database, Genome Information Broker (GIB), has been improved by incorporating XML. It is now possible to perform a more sophisticated database search against the new GIB than the ordinary BLAST or FASTA search.

Animals↗