Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

A comparison of physical mapping algorithms based on the maximum likelihood model.

MOTIVATION: Physical mapping of chromosomes using the maximum likelihood (ML) model is a problem of high computational complexity entailing both discrete optimization to recover the optimal probe order as well as continuous optimization to recover the optimal inter-probe spacings. In this paper, two versions of the genetic algorithm (GA) are proposed, one with heuristic crossover and deterministic replacement and the other with heuristic crossover and stochastic replacement, for the physical mapping problem under the maximum likelihood model. The genetic algorithms are compared with two other discrete optimization approaches, namely simulated annealing (SA) and large-step Markov chains (LSMC), in terms of solution quality and runtime efficiency. RESULTS: The physical mapping algorithms based on the GA, SA and LSMC have been tested using synthetic datasets and real datasets derived from cosmid libraries of the fungus Neurospora crassa. The GA, especially the version with heuristic crossover and stochastic replacement, is shown to consistently outperform the SA-based and LSMC-based physical mapping algorithms in terms of runtime and final solution quality. Experimental results on real datasets and simulated datasets are presented. Further improvements to the GA in the context of physical mapping under the maximum likelihood model are proposed. AVAILABILITY: The software is available upon request from the first author.

Algorithms↗

Identification of susceptibility loci for complex diseases in a case-control association study using the Genetic Analysis Workshop 14 dataset.

Although current methods in genetic epidemiology have been extremely successful in identifying genetic loci responsible for Mendelian traits, most common diseases do not follow simple Mendelian modes of inheritance. It is important to consider how our current methodologies function in the realm of complex diseases. The aim of this study was to determine the ability of conventional association methods to fine map a locus of interest. Six study populations were selected from 10 replicates (New York) from the Genetic Analysis Workshop 14 simulated dataset and analyzed for association between the disease trait and locus D2. Genotypes from 45 single-nucleotide polymorphisms in the telomeric region of chromosome 3 were analyzed by Pearson's chi-square tests for independence to test for association with the disease trait of interest. A significant association was detected within the region; however, it was found 3 cM from the documented location of the D2 disease locus. This result was most likely due to the method used for data simulation. In general, this study showed that conventional case-control association methods could detect disease loci responsible for the development of complex traits.

Case-Control Studies↗

A practical guide to mitochondrial DNA error prevention in clinical, forensic, and population genetics.

Several suggestions have been made for avoiding errors in mitochondrial DNA (mtDNA) sequencing and documentation. Unfortunately, the current clinical, forensic, and population genetic literature on mtDNA still delivers a large number of studies with flawed sequence data, which, in extreme cases, damage the whole message of a study. The phylogenetic approach has been shown to be useful for pinpointing most of the errors. However, many geneticists, especially in the forensic and medical fields, are not familiar with either effective search strategies or the evolutionary terminology. We here provide a manual that should help prevent errors at any stage by re-examining data fresh from the sequencer in the light of previously published data. A fictitious case study of a European mtDNA data set (albeit composed from the literature) then demonstrates the steps one has to go through in order to assess the quality of sequencing and documentation.

DNA, Mitochondrial↗

A flexible data analysis tool for chemical genetic screens.

High-throughput assays generate immense quantities of data that require sophisticated data analysis tools. We have created a freely available software tool, SLIMS (Small Laboratory Information Management System), for chemical genetics which facilitates the collection and analysis of large-scale chemical screening data. Compound structures, physical locations, and raw data can be loaded into SLIMS. Raw data from high-throughput assays are normalized using flexible analysis protocols, and systematic spatial errors are automatically identified and corrected. Various computational analyses are performed on tested compounds, and dilution-series data are processed using standard or user-defined algorithms. Finally, published literature associated with active compounds is automatically retrieved from Medline and processed to yield potential mechanisms of actions. SLIMS provides a framework for analyzing high-throughput assay data both as a laboratory information management system and as a platform for experimental analysis.

Cyclic AMP Response Element-Binding Protein↗

Genetic diversity of Crotalaria germplasm assessed through phylogenetic analysis of EST-SSR markers.

The genetic diversity of the genus Crotalaria is unknown even though many species in this genus are economically valuable. We report the first study in which polymorphic expressed sequence tag-simple sequence repeat (EST-SSR) markers derived from Medicago and soybean were used to assess the genetic diversity of the Crotalaria germplasm collection. This collection consisted of 26 accessions representing 4 morphologically characterized species. Phylogenetic analysis partitioned accessions into 4 main groups generally along species lines and revealed that 2 accessions were incorrectly identified as Crotalaria juncea and Crotalaria spectabilis instead of Crotalaria retusa. Morphological re-examination confirmed that these 2 accessions were misclassified during curation or conservation and were indeed C. retusa. Some amplicons from Crotalaria were sequenced and their sequences showed a high similarity (89% sequence identity) to Medicago truncatula from which the EST-SSR primers were designed; however, the SSRs were completely deleted in Crotalaria. Highly distinguishing markers or more sequences are required to further classify accessions within C. juncea.

Base Sequence↗

The distribution of SNPs in human gene regulatory regions.

BACKGROUND: As a result of high-throughput genotyping methods, millions of human genetic variants have been reported in recent years. To efficiently identify those with significant biological functions, a practical strategy is to concentrate on variants located in important sequence regions such as gene regulatory regions. RESULTS: Analysis of the most common type of variant, single nucleotide polymorphisms (SNPs), shows that in gene promoter regions more SNPs occur in close proximity to transcriptional start sites than in regions further upstream, and a disproportionate number of those SNPs represent nucleotide transversions. Additionally, the number of SNPs found in the predicted transcription factor binding sites is higher than in non-binding site sequences. CONCLUSION: Current information about transcription factor binding site sequence patterns may not be exhaustive, and SNPs may be actively involved in influencing gene expression by affecting the transcription factor binding sites.

Binding Sites↗

Functional partitioning of yeast co-expression networks after genome duplication.

Several species of yeast, including the baker's yeast Saccharomyces cerevisiae, underwent a genome duplication roughly 100 million years ago. We analyze genetic networks whose members were involved in this duplication. Many networks show detectable redundancy and strong asymmetry in their interactions. For networks of co-expressed genes, we find evidence for network partitioning whereby the paralogs appear to have formed two relatively independent subnetworks from the ancestral network. We simulate the degeneration of networks after duplication and find that a model wherein the rate of interaction loss depends on the "neighborliness" of the interacting genes produces networks with parameters similar to those seen in the real partitioned networks. We propose that the rationalization of network structure through the loss of pair-wise gene interactions after genome duplication provides a mechanism for the creation of semi-independent daughter networks through the division of ancestral functions between these daughter networks.

Base Sequence↗

Genetic research on the UK population--do new principles need to be developed?

Governments around the world are beginning to generate population databases as resources for genetic research. In the UK, a proposal has been tabled that plans to incorporate National Health Service information--a move that will effectively create a database of around 60 million. However, this new population collection will not conform to standards established by other national genetic databases, and the UK government report has not accounted for key ethical issues.

Bioethics↗

Predicting risk of coronary artery disease from DNA microarray-based genotyping using neural networks and other statistical analysis tool.

This paper presents a novel approach for complex disease prediction that we have developed, exemplified by a study on risk of coronary artery disease (CAD). This multi-disciplinary approach straddles fields of microarray technology and genetics, neural networks (NN), data mining and machine learning, as well as traditional statistical analysis techniques, namely principal components analysis (PCA) and factor analysis (FA). A description of the biological background of the study is given, followed by a detailed description of how the problem has been modeled for analyses by neural networks and FA. A committee learning approach for NN has been used to improve generalization rates. We show that our NN approach is able to yield promising prediction results despite using only the most fundamental network structures. More interestingly, through the statistical analysis process, genes of similar biological functions have been clustered. In addition, a gene marker involved in breaking down lipids has been found to be the most correlated to CAD.

Algorithms↗

Genes, age, and alcoholism: analysis of GAW14 data.

A genetic analysis of age of onset of alcoholism was performed on the Collaborative Study on the Genetics of Alcoholism data released for Genetic Analysis Workshop 14. Our study illustrates an application of the log-normal age of onset model in our software Genetic Epidemiology Models (GEMs). The phenotype ALDX1 of alcoholism was studied. The analysis strategy was to first find the markers of the Affymetrix SNP dataset with significant association with age of onset, and then to perform linkage analysis on them. ALDX1 revealed strong evidence of linkage for marker tsc0041591 on chromosome 2 and suggestive linkage for marker tsc0894042 on chromosome 3. The largest separation in mean ages of onset of ALDX1 was 19.76 and 24.41 between male smokers who are carriers of the risk allele of tsc0041591 and the non-carriers, respectively. Hence, male smokers who are carriers of marker tsc0041591 on chromosome 2 have an average onset of ALDX1 almost 5 years earlier than non-carriers.

Age of Onset↗

Origin, genetic diversity, and population structure of Chinese domestic sheep.

To characterize the origin, genetic diversity, and phylogeographic structure of Chinese domestic sheep, we here analyzed a 531-bp fragment of mtDNA control region of 449 Chinese autochthonous sheep from 19 breeds/populations from 13 geographic regions, together with previously reported 44 sequences from Chinese indigenous sheep. Phylogenetic analysis showed that all three previously defined lineages A, B, and C were found in all sampled Chinese sheep populations, except for the absence of lineage C in four populations. Network profiles revealed that the lineages B and C displayed a star-like phylogeny with the founder haplotype in the centre, and that two star-like subclades with two founder haplotypes were identified in lineage A. The pattern of genetic variation in lineage A, together with the divergence time between the two central founder haplotypes suggested that two independent domestication events have occurred in sheep lineage A. Considerable mitochondrial diversity was observed in Chinese sheep. Weak structuring was observed either among Chinese indigenous sheep populations or between Asian and European sheep and this can be attributable to long-term strong gene flow induced by historical human movements. The high levels of intra-population diversity in Chinese sheep and the weak phylogeographic structuring indicated three geographically independent domestication events have occurred and the domestication place was not only confined to the Near East, but also occurred in other regions.

Animal Migration↗

Quantitative trait locus-specific genotype x alcoholism interaction on linkage for evoked electroencephalogram oscillations.

We explored the evidence for a quantitative trait locus (QTL)-specific genotype x alcoholism interaction for an evoked electroencephalogram theta band oscillation (ERP) phenotype on a region of chromosome 7 in participants of the US Collaborative Study on the Genetics of Alcoholism. Among 901 participants with both genotype and phenotype data available, we performed variance component linkage analysis (SOLAR version 2.1.2) in the full sample and stratified by DSM-III-R and Feighner-definite alcoholism categories. The heritability of the ERP phenotype after adjusting for age and sex effects in the combined sample and in the alcoholism classification sub-groups ranged from 40% to 66%. Linkage on chromosome 7 was identified at 158 cM (LOD = 3.8) in the full sample and at 108 in the non-alcoholic subgroup (LOD = 3.1). Further, we detected QTL-specific genotype x alcoholism interaction at these loci. This work demonstrates the importance of considering the complexity of common complex traits in our search for genes that predispose to alcoholism.

Alcoholism↗

Intergenic mRNAs. Minor gene products or tools of diversity?

The general model of gene expression implies that the units of genetic information, the genes, constitute the basis of the mRNA and protein products detected in living organisms. It is well known that in eukaryotic cells a single gene may give rise to various gene products due to alternative splicing and/or different positions of transcriptional start and termination. However, recent experimental results suggest that in addition to this variation in gene expression, another level of complexity may also exist as evidence for the presence of mRNAs that combine sequence information (exons) from distinct genes is accumulating. Moreover, the mechanisms that allow this production of intergenic mRNAs are starting to unravel and it appears that these follow two general pathways: (a) bypass of transcriptional termination, resulting in the generation of bicistronic pre-mRNAs; and (b) authentic trans-splicing events between pre-mRNAs of distinct genes.

Animals↗

Functional mapping and annotation of genetic associations with FUMA.

A main challenge in genome-wide association studies (GWAS) is to pinpoint possible causal variants. Results from GWAS typically do not directly translate into causal variants because the majority of hits are in non-coding or intergenic regions, and the presence of linkage disequilibrium leads to effects being statistically spread out across multiple variants. Post-GWAS annotation facilitates the selection of most likely causal variant(s). Multiple resources are available for post-GWAS annotation, yet these can be time consuming and do not provide integrated visual aids for data interpretation. We, therefore, develop FUMA: an integrative web-based platform using information from multiple biological resources to facilitate functional annotation of GWAS results, gene prioritization and interactive visualization. FUMA accommodates positional, expression quantitative trait loci (eQTL) and chromatin interaction mappings, and provides gene-based, pathway and tissue enrichment results. FUMA results directly aid in generating hypotheses that are testable in functional experiments aimed at proving causal relations.

Chromatin↗

Exploitation of colinear relationships between the genomes of Lotus japonicus, Pisum sativum and Arabidopsis thaliana, for positional cloning of a legume symbiosis gene.

The Lotus japonicus LjSYM2 gene, and the Pisum sativum orthologue PsSYM19, are required for the formation of nitrogen-fixing root nodules and arbuscular mycorrhiza. Here we describe the map-based cloning procedure leading to the isolation of both genes. Marker information from a classical AFLP marker-screen in Lotus was integrated with a comparative genomics approach, utilizing Arabidopsis genome sequence information and the pea genetic map. A network of gene-based markers linked in all three species was identified, suggesting local colinearity in the region around LjSYM2/PsSYM19. The closest AFLP marker was located just over 200 kb from the LjSYM2 gene, the marker SHMT, which was converted from a marker on the pea map, was only 7.9 kb away. The LjSYM2/PsSYM19 region corresponds to two duplicated segments of the Arabidopsis chromosomes AtII and AtIV. Lotus homologues of Arabidopsis genes within these segments were mapped to three clusters on LjI, LjII and LjVI, suggesting that during evolution the genomic segment surrounding LjSYM2 has been subjected to duplication events. However, one marker, AUX-1, was identified based on colinearity between Lotus and Arabidopsis that mapped in physical proximity of the LjSym2 gene.

Arabidopsis↗

Detecting disease-causing mutations in the human genome by haplotype matching.

Comparisons between haplotypes from affected patients and the human reference genome are frequently used to identify candidates for disease-causing mutations, even though these alignments are expected to reveal a high level of background neutral polymorphism. This limits the scope of genetic studies to relatively small genomic intervals, because current methods for distinguishing potential causal mutations from neutral variation are inefficient. Here we describe a new strategy for detecting mutations that is based on comparing affected haplotypes with closely matched control sequences from healthy individuals, rather than with the human reference genome. We use theory, simulation, and a real data set to show that this approach is expected to reduce the number of sequence variants that must be subjected to follow-up analysis by at least a factor of 20 when closely matched control sequences are selected from a reference panel with as few as 100 control genomes. We also define a reference data resource that would allow efficient application of this strategy to large critical intervals across the genome.

Alleles↗

Feature selection for splice site prediction: a new method using EDA-based feature ranking.

BACKGROUND: The identification of relevant biological features in large and complex datasets is an important step towards gaining insight in the processes underlying the data. Other advantages of feature selection include the ability of the classification system to attain good or even better solutions using a restricted subset of features, and a faster classification. Thus, robust methods for fast feature selection are of key importance in extracting knowledge from complex biological data. RESULTS: In this paper we present a novel method for feature subset selection applied to splice site prediction, based on estimation of distribution algorithms, a more general framework of genetic algorithms. From the estimated distribution of the algorithm, a feature ranking is derived. Afterwards this ranking is used to iteratively discard features. We apply this technique to the problem of splice site prediction, and show how it can be used to gain insight into the underlying biological process of splicing. CONCLUSION: We show that this technique proves to be more robust than the traditional use of estimation of distribution algorithms for feature selection: instead of returning a single best subset of features (as they normally do) this method provides a dynamical view of the feature selection process, like the traditional sequential wrapper methods. However, the method is faster than the traditional techniques, and scales better to datasets described by a large number of features.

Adenosine↗

Description of the data from the Collaborative Study on the Genetics of Alcoholism (COGA) and single-nucleotide polymorphism genotyping for Genetic Analysis Workshop 14.

The data provided to the Genetic Analysis Workshop 14 (GAW 14) was the result of a collaboration among several different groups, catalyzed by Elizabeth Pugh from The Center for Inherited Disease Research (CIDR) and the organizers of GAW 14, Jean MacCluer and Laura Almasy. The DNA, phenotypic characterization, and microsatellite genomic survey were provided by the Collaborative Study on the Genetics of Alcoholism (COGA), a nine-site national collaboration funded by the National Institute of Alcohol and Alcoholism (NIAAA) and the National Institute of Drug Abuse (NIDA) with the overarching goal of identifying and characterizing genes that affect the susceptibility to develop alcohol dependence and related phenotypes. CIDR, Affymetrix, and Illumina provided single-nucleotide polymorphism genotyping of a large subset of the COGA subjects. This article briefly describes the dataset that was provided.

Alcoholism↗