Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Toxicology and genetic toxicology in the new era of "toxicogenomics": impact of "-omics" technologies.

The unprecedented advances in molecular biology during the last two decades have resulted in a dramatic increase in knowledge about gene structure and function, an immense database of genetic sequence information, and an impressive set of efficient new technologies for monitoring genetic sequences, genetic variation, and global functional gene expression. These advances have led to a new sub-discipline of toxicology: "toxicogenomics". We define toxicogenomics as "the study of the relationship between the structure and activity of the genome (the cellular complement of genes) and the adverse biological effects of exogenous agents". This broad definition encompasses most of the variations in the current usage of this term, and in its broadest sense includes studies of the cellular products controlled by the genome (messenger RNAs, proteins, metabolites, etc.). The new "global" methods of measuring families of cellular molecules, such as RNA, proteins, and intermediary metabolites have been termed "-omic" technologies, based on their ability to characterize all, or most, members of a family of molecules in a single analysis. With these new tools, we can now obtain complete assessments of the functional activity of biochemical pathways, and of the structural genetic (sequence) differences among individuals and species, that were previously unattainable. These powerful new methods of high-throughput and multi-endpoint analysis include gene expression arrays that will soon permit the simultaneous measurement of the expression of all human genes on a single "chip". Likewise, there are powerful new methods for protein analysis (proteomics: the study of the complement of proteins in the cell) and for analysis of cellular small molecules (metabonomics: the study of the cellular metabolites formed and degraded under genetic control). This will likely be extended in the near future to other important classes of biomolecules such as lipids, carbohydrates, etc. These assays provide a general capability for global assessment of many classes of cellular molecules, providing new approaches to assessing functional cellular alterations. These new methods have already facilitated significant advances in our understanding of the molecular responses to cell and tissue damage, and of perturbations in functional cellular systems. As a result of this rapidly changing scientific environment, regulatory and industrial toxicology practice is poised to undergo dramatic change during the next decade. These advances present exciting opportunities for improved methods of identifying and evaluating potential human and environmental toxicants, and of monitoring the effects of exposures to these toxicants. These advances also present distinct challenges. For example, the significance of specific changes and the performance characteristics of new methods must be fully understood to avoid misinterpretation of data that could lead to inappropriate conclusions about the toxicity of a chemical or a mechanism of action. We discuss the likely impact of these advances on the fields of general and genetic toxicology, and risk assessment. We anticipate that these new technologies will (1) lead to new families of biomarkers that permit characterization and efficient monitoring of cellular perturbations, (2) provide an increased understanding of the influence of genetic variation on toxicological outcomes, and (3) allow definition of environmental causes of genetic alterations and their relationship to human disease. The broad application of these new approaches will likely erase the current distinctions among the fields of toxicology, pathology, genetic toxicology, and molecular genetics. Instead, a new integrated approach will likely emerge that involves a comprehensive understanding of genetic control of cellular functions, and of cellular responses to alterations in normal molecular structure and function.

Animals↗

A genetic regulatory network for Xenopus mesendoderm formation.

We have constructed a genetic regulatory network (GRN) summarising the functional relationships between the transcription factors (TFs) and embryonic signals involved in Xenopus mesendoderm formation. It is supported by a relational database containing the experimental evidence and both are available in interactive form via the World Wide Web. This network highlights areas for further study and provides a framework for systematic interrogation of new data. Comparison with the equivalent network for the sea urchin identifies conserved features of the deuterostome ancestral pathway, including positive feedback loops, GATA factors, SoxB, Brachyury and a previously underemphasised role for beta-catenin. In contrast, some features central to one species have not yet been found in the other, for example, Krox and Otx in sea urchin, and Mix and Nodal in Xenopus. Such differences may represent evolved features or may eventually be resolved. For example, in Xenopus, Nodal-related genes are positively regulated by beta-catenin and at least one of them is repressed by Sox3, as is the uncharacterised early signal (ES) inducing endomesoderm in the sea urchin, suggesting that ES may be a Nodal-like TGF-beta. Wider comparisons of such networks will inform our understanding of developmental evolution.

Animals↗

Identification of proteins from non-model organisms using mass spectrometry: application to a hibernating mammal.

A major challenge in the life sciences is the extraction of detailed molecular information from plants and animals that are not among the handful of exhaustively studied "model organisms." As a consequence, certain species with novel phenotypes are often ignored due to the lack of searchable databases, tractable genetics, stock centers, and more recently, a sequenced genome. Characterization of phenotype at the molecular level commonly relies on the identification of differentially expressed proteins by combining database searching with tandem mass spectrometry (MS) of peptides derived from protein fragmentation. However, the identification of short peptides from nonmodel organisms can be hampered by the lack of sufficient amino acid sequence homology with proteins in existing databases; therefore, a database search strategy that encompasses both identity and homology can provide stronger evidence than a single search alone. The use of multiple algorithms for database searches may also increase the probability of correct protein identification since it is unlikely that each program would produce false negative or positive hits for the same peptides. In this study, four software packages, Mascot, Pro ID, Sequest, and Pro BLAST, were compared in their ability to identify proteins from the thirteen-lined ground squirrel (Spermophilus tridecemlineatus), a hibernating mammal that lacks a completely sequenced genome. Our results show similarities as well as the degree of variability among different software packages when the identical protein database is searched. In the process of this study, we identified the up-regulation of succinyl CoA-transferase (SCOT) in the heart of hibernators. SCOT is the rate-limiting enzyme in the catabolism of ketone bodies, an important alternative fuel source during hibernation.

Algorithms↗

Merging protein, gene and genomic data: the evolution of the MDR-ADH family.

Multiple members of the MDR-ADH (MDR: Medium-chain dehydrogenases/reductases; ADH: alcohol dehydrogenase) family are found in vertebrates, although the enzymes that belong to this family have also been isolated from bacteria, yeast, plant and animal sources. Initial understanding of the physiological roles and evolution of the family relied on biochemical studies, protein alignments and protein structure comparisons. Subsequently, studies at the genetic level yielded new information: the expression pattern, exon-intron distribution, in silico-derived protein sequences and murine knockout phenotypes. More recently, genomic and EST databases have revealed new family members and the chromosomal location and position in the cluster of both the first and new forms. The data now available provide a comprehensive scenario, from which a reliable picture of the evolutionary history of this family can be made.

Alcohol Dehydrogenase↗

An ordered, nonredundant library of Pseudomonas aeruginosa strain PA14 transposon insertion mutants.

Random transposon insertion libraries have proven invaluable in studying bacterial genomes. Libraries that approach saturation must be large, with multiple insertions per gene, making comprehensive genome-wide scanning difficult. To facilitate genome-scale study of the opportunistic human pathogen Pseudomonas aeruginosa strain PA14, we constructed a nonredundant library of PA14 transposon mutants (the PA14NR Set) in which nonessential PA14 genes are represented by a single transposon insertion chosen from a comprehensive library of insertion mutants. The parental library of PA14 transposon insertion mutants was generated by using MAR2xT7, a transposon compatible with transposon-site hybridization and based on mariner. The transposon-site hybridization genetic footprinting feature broadens the utility of the library by allowing pooled MAR2xT7 mutants to be individually tracked under different experimental conditions. A public, internet-accessible database (the PA14 Transposon Insertion Mutant Database, http://ausubellab.mgh.harvard.edu/cgi-bin/pa14/home.cgi) was developed to facilitate construction, distribution, and use of the PA14NR Set. The usefulness of the PA14NR Set in genome-wide scanning for phenotypic mutants was validated in a screen for attachment to abiotic surfaces. Comparison of the genes disrupted in the PA14 transposon insertion library with an independently constructed insertion library in P. aeruginosa strain PAO1 provides an estimate of the number of P. aeruginosa essential genes.

DNA Transposable Elements↗

Gene mining: a novel and powerful ensemble decision approach to hunting for disease genes using microarray expression profiling.

Current applications of microarrays focus on precise classification or discovery of biological types, for example tumor versus normal phenotypes in cancer research. Several challenging scientific tasks in the post-genomic epoch, like hunting for the genes underlying complex diseases from genome-wide gene expression profiles and thereby building the corresponding gene networks, are largely overlooked because of the lack of an efficient analysis approach. We have thus developed an innovative ensemble decision approach, which can efficiently perform multiple gene mining tasks. An application of this approach to analyze two publicly available data sets (colon data and leukemia data) identified 20 highly significant colon cancer genes and 23 highly significant molecular signatures for refining the acute leukemia phenotype, most of which have been verified either by biological experiments or by alternative analysis approaches. Furthermore, the globally optimal gene subsets identified by the novel approach have so far achieved the highest accuracy for classification of colon cancer tissue types. Establishment of this analysis strategy has offered the promise of advancing microarray technology as a means of deciphering the involved genetic complexities of complex diseases.

Acute Disease↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

Improving literature based discovery support by genetic knowledge integration.

We present an interactive literature based biomedical discovery support system (BITOLA). The goal of the system is to discover new, potentially meaningful relations between a given starting concept of interest and other concepts, by mining the bibliographic database Medline. To make the system more suitable for disease candidate gene discovery and to decrease the number of candidate relations, we integrate background knowledge about the chromosomal location of the starting disease as well as the chromosomal location of the candidate genes from resources such as LocusLink, HUGO and OMIM. The BITOLA system can be also used as an alternative way of searching the Medline database. The system is available at http://www.mf.uni-lj.si/bitola/.

Algorithms↗

Construction of two genetic linkage maps in cultivated tetraploid alfalfa (Medicago sativa) using microsatellite and AFLP markers.

BACKGROUND: Alfalfa (Medicago sativa) is a major forage crop. The genetic progress is slow in this legume species because of its autotetraploidy and allogamy. The genetic structure of this species makes the construction of genetic maps difficult. To reach this objective, and to be able to detect QTLs in segregating populations, we used the available codominant microsatellite markers (SSRs), most of them identified in the model legume Medicago truncatula from EST database. A genetic map was constructed with AFLP and SSR markers using specific mapping procedures for autotetraploids. The tetrasomic inheritance was analysed in an alfalfa mapping population. RESULTS: We have demonstrated that 80% of primer pairs defined on each side of SSR motifs in M. truncatula EST database amplify with the alfalfa DNA. Using a F1 mapping population of 168 individuals produced from the cross of 2 heterozygous parental plants from Magali and Mercedes cultivars, we obtained 599 AFLP markers and 107 SSR loci. All but 3 SSR loci showed a clear tetrasomic inheritance. For most of the SSR loci, the double-reduction was not significant. For the other loci no specific genotypes were produced, so the significant double-reduction could arise from segregation distortion. For each parent, the genetic map contained 8 groups of four homologous chromosomes. The lengths of the maps were 2649 and 3045 cM, with an average distance of 7.6 and 9.0 cM between markers, for Magali and Mercedes parents, respectively. Using only the SSR markers, we built a composite map covering 709 cM. CONCLUSIONS: Compared to diploid alfalfa genetic maps, our maps cover about 88-100% of the genome and are close to saturation. The inheritance of the codominant markers (SSR) and the pattern of linkage repulsions between markers within each homology group are consistent with the hypothesis of a tetrasomic meiosis in alfalfa. Except for 2 out of 107 SSR markers, we found a similar order of markers on the chromosomes between the tetraploid alfalfa and M. truncatula genomes indicating a high level of colinearity between these two species. These maps will be a valuable tool for alfalfa breeding and are being used to locate QTLs.

Alleles↗

First sequenced mitochondrial genome from the phylum Acanthocephala (Leptorhynchoides thecatus) and its phylogenetic position within Metazoa.

The complete sequence of the mitochondrial genome of Leptorhynchoides thecatus (Acanthocephala) was determined, and a phylogenetic analysis was carried out to determine its placement within Metazoa. The genome is circular, 13,888 bp, and contains at least 36 of the 37 genes typically found in animal mitochondrial genomes. The genes for the large and small ribosomal RNA subunits are shorter than those of most metazoans, and the structures of most of the tRNA genes are atypical. There are two significant noncoding regions (377 and 294 bp), which are the best candidates for a control region; however, these regions do not appear similar to any of the control regions of other animals studied to date. The amino acid and nucleotide sequences of the protein coding genes of L. thecatus and 25 other metazoan taxa were used in both maximum likelihood and maximum parsimony phylogenetic analyses. Results indicate that among taxa with available mitochondrial genome sequences, Platyhelminthes is the closest relative to L. thecatus, which together are the sister taxon of Nematoda; however, long branches and/or base composition bias could be responsible for this result. The monophyly of Ecdysozoa, molting organisms, was not supported by any of the analyses. This study represents the first mitochondrial genome of an acanthocephalan to be sequenced and will allow further studies of systematics, population genetics, and genome evolution.

Acanthocephala↗

A Bayesian approach for constructing genetic maps when markers are miscoded.

The advent of molecular markers has created opportunities for a better understanding of quantitative inheritance and for developing novel strategies for genetic improvement of agricultural species, using information on quantitative trait loci (QTL). A QTL analysis relies on accurate genetic marker maps. At present, most statistical methods used for map construction ignore the fact that molecular data may be read with error. Often, however, there is ambiguity about some marker genotypes. A Bayesian MCMC approach for inferences about a genetic marker map when random miscoding of genotypes occurs is presented, and simulated and real data sets are analyzed. The results suggest that unless there is strong reason to believe that genotypes are ascertained without error, the proposed approach provides more reliable inference on the genetic map.

Bayes Theorem↗

iVici: Interrelational Visualization and Correlation Interface.

We have developed an application, iVici, to analyze cellular networks represented as addressable symmetric or asymmetric two-dimensional matrices. iVici was designed to permit simultaneous visualization and correlation of multiple datasets, representing any relationship between a set of genes, mRNAs, or proteins. Visual overlay of datasets and addressable access to gene annotations permits comparison of networks of different types (for example protein-protein interactions and genetic networks) or investigation of the dynamic reorganization of a particular network.

Computational Biology↗

Single nucleotide polymorphisms associated with rat expressed sequences.

Single nucleotide polymorphisms (SNPs) are the most common source of genetic variation in populations and are thus most likely to account for the majority of phenotypic and behavioral differences between individuals or strains. Although the rat is extensively studied for the latter, data on naturally occurring polymorphisms are mostly lacking. We have used publicly available sequences consisting of whole-genome shotgun (WGS), expressed sequence tag (EST), and mRNA data as a source for the in silico identification of SNPs in gene-coding regions and have identified a large collection of 33,305 high-quality candidate SNPs. Experimental verification of 471 candidate SNPs using a limited set of rat isolates revealed a confirmation rate of approximately 50%. Although the majority of SNPs were identified between Sprague-Dawley (EST data) and Brown Norway (WGS data) strains, we found that 66% of the verified variations are common among different rat strains. All SNPs were extensively annotated, including chromosomal and genetic map information, and nonsynonymous SNPs were analyzed by SIFT and PolyPhen prediction programs for their potential deleterious effect on protein function. Interestingly, we retrieved three SNPs from the database that result in the introduction of a premature stop codon and that could be confirmed experimentally. Two of these "in silico-identified knockouts" reside in interesting QTL regions. Data are publicly available via a Web interface (http://cascad.niob.knaw.nl), allowing simple and advanced search queries.

Animals↗

The importance of ignorance.

The need for people to keep their genetic data confidential is crucial to help exploit medical advances, a key British Committee believes. Nigel Williams reports.

Databases, Genetic↗

Survey of polymorphic sequence variation in the immediate 5' region of human DNA repair genes.

Systematic screens have revealed extensive DNA sequence variation existing in the human population. Studies of the role of polymorphic genetic variants in explaining the association of family history with risk of common disease have generally focused on variants predicted to disrupt protein structure and activity. Recent studies have identified genetic variation in the level of expression of many genes, variation that is potentially biologically relevant in explaining individual variation in disease risk. In a survey of data available for 108 DNA repair genes that have been systematically screened for sequence variation, an average of 3.3 SNPs per gene were found to exist at a variant allele frequency of at least 0.02 in the region 2kb upstream from the 5'-untranslated region. One-third of the genes harbored a SNP with an allele frequency of at least 0.02 within a predicted promotor element. These variants are distributed among promoter elements that average 20 elements per gene. The frequency of polymorphic SNPs in CpG islands was 0.8 per gene, while the frequency of SNPs in the 5'-UTR was 0.7 per gene. The recognition of extensive genetic variation with potential to impact levels of gene expression, and thereby exacerbate the impact of amino acid substitution variants on the activity of proteins, increases the complexity of analyses required to explain the molecular genetic basis for the familial contribution to the sporadic incidence of common disease.

5' Untranslated Regions↗