Genetic Analysis Workshop 14: microsatellite and single-nucleotide polymorphism marker loci for genome-wide scans.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In this paper we will examine some ethical aspects of the role that computers and computing increasingly play in new genetics. Our claim is that there is no new genetics without computer science. Computer science is important for the new genetics on two levels: (1) from a theoretical perspective, and (2) from the point of view of geneticists practice. With respect to (1), the new genetics is fully impregnate with concepts that are basic for computer science. Regarding (2), recent developments in the Human Genome Project (HGP) have shown that computers shape the practices of molecular genetics; an important example is the Shotgun Method's contribution to accelerating the mapping of the human genome. A new challenge to the HGP is provided by the Open Source Philosophy (I computer science), which is another way computer technologies now influence the shaping of public policy debates involving genomics.
The prediction of translation initiation sites (TISs) in eukaryotic mRNAs has been a challenging problem in computational molecular biology. In this paper, we present a new algorithm to recognize TISs with a very high accuracy. Our algorithm includes two novel ideas. First, we introduce a class of new sequence-similarity kernels based on string editing, called edit kernels, for use with support vector machines (SVMs) in a discriminative approach to predict TISs. The edit kernels are simple and have significant biological and probabilistic interpretations. Although the edit kernels are not positive definite, it is easy to make the kernel matrix positive definite by adjusting the parameters. Second, we convert the region of an input mRNA sequence downstream to a putative TIS into an amino acid sequence before applying SVMs to avoid the high redundancy in the genetic code. The algorithm has been implemented and tested on previously published data. Our experimental results on real mRNA data show that both ideas improve the prediction accuracy greatly and that our method performs significantly better than those based on neural networks and SVMs with polynomial kernels or Salzberg kernels.
Multivariate linkage analysis using several correlated traits may provide greater statistical power to detect susceptibility genes in loci whose effects are too small to be detected in univariate analysis. In this analysis, we apply a new approach and perform a linkage analysis of several electrophysiological phenotypes of the Collaborative Study on the Genetics of Alcoholism data of the Genetic Analysis Workshop 14. Our approach is based on a variance-component model to map candidate genes using repeated or longitudinal measurements. It can take into account covariate effects and time-dependent genetic effects in general pedigree data. We compare our results with the ones obtained by SOLAR using single measurement data. Our multivariate linkage analysis found linkage evidence on two regions on chromosome 4: around marker GABRB1 at 51.4 cM and marker FABP2 at 116.8 cM (unadjusted p-value = 0.00006).
We compared linkage analysis results for an alcoholism trait, ALDX1 (DSM-III-R and Feigner criteria) using a nonparametric linkage analysis method, which takes into account allele sharing among several affected persons, for both microsatellite and single-nucleotide polymorphism (SNP) markers (Affymetrix and Illumina) in the Collaborative Study on the Genetics of Alcoholism (COGA) dataset provided to participants at the Genetic Analysis Workshop 14 (GAW14). The two sets of linkage results from the dense Affymetrix SNP markers and less densely spaced Illumina SNP markers are very similar. The linkage analysis results from microsatellite and SNP markers are generally similar, but the match is not perfect. Strong linkage peaks were found on chromosome 7 in three sets of linkage analyses using both SNP and microsatellite marker data. We also observed that for SNP markers, using the given genetic map and using the map by converting 1 megabase pair (1 Mb) to 1 centimorgan (cM), did not change the linkage results. We recommend the use of the 1 Mb-to-1 cM converted map in a first round of linkage analysis with SNP markers in which map integration is an issue.
BACKGROUND: The cellular response of plants to water-deficits has both economic and evolutionary importance directly affecting plant productivity in agriculture and plant survival in the natural environment. Genes induced by water-deficit stress have been successfully enumerated in plants that are relatively sensitive to cellular dehydration, however we have little knowledge as to the adaptive role of these genes in establishing tolerance to water loss at the cellular level. Our approach to address this problem has been to investigate the genetic responses of plants that are capable of tolerating extremes of dehydration, in particular the desiccation-tolerant bryophyte, Tortula ruralis. To establish a sound basis for characterizing the Tortula genome in regards to desiccation tolerance, we analyzed 10,368 expressed sequence tags (ESTs) from rehydrated rapid-dried Tortula gametophytes, a stage previously determined to exhibit the maximum stress induced change in gene expression. RESULTS: The 10, 368 ESTs formed 5,563 EST clusters (contig groups representing individual genes) of which 3,321 (59.7%) exhibited similarity to genes present in the public databases and 2,242 were categorized as unknowns based on protein homology scores. The 3,321 clusters were classified by function using the Gene Ontology (GO) hierarchy and the KEGG database. The results indicate that the transcriptome contains a diverse population of transcripts that reflects, as expected, a period of metabolic upheaval in the gametophyte cells. Much of the emphasis within the transcriptome is centered on the protein synthetic machinery, ion and metabolite transport, and membrane biosynthesis and repair. Rehydrating gametophytes also have an abundance of transcripts that code for enzymes involved in oxidative stress metabolism and phosphorylating activities. The functional classifications reflect a remarkable consistency with what we have previously established with regards to the metabolic activities that are important in the recovery of the gametophytes from desiccation. A comparison of the GO distribution of Tortula clusters with an identical analysis of 9,981 clusters from the desiccation sensitive bryophyte species Physcomitrella patens, revealed, and accentuated, the differences between stressed and unstressed transcriptomes. Cross species sequence comparisons indicated that on the whole the Tortula clusters were more closely related to those from Physcomitrella than Arabidopsis (complete genome BLASTx comparison) although because of the differences in the databases there were more high scoring matches to the Arabidopsis sequences. The most abundant transcripts contained within the Tortula ESTs encode Late Embryogenesis Abundant (LEA) proteins that are normally associated with drying plant tissues. This suggests that LEAs may also play a role in recovery from desiccation when water is reintroduced into a dried tissue. CONCLUSION: The establishment of a rehydration EST collection for Tortula ruralis, an important plant model for plant stress responses and vegetative desiccation tolerance, is an important step in understanding the genome level response to cellular dehydration. The type of transcript analysis performed here has laid the foundation for more detailed functional and genome level analyses of the genes involved in desiccation tolerance in plants.
Ontologies are useful for organizing large numbers of concepts having complex relationships, such as the breadth of genetic and clinical knowledge in pharmacogenomics. But because ontologies change and knowledge evolves, it is time consuming to maintain stable mappings to external data sources that are in relational format. We propose a method for interfacing ontology models with data acquisition from external relational data sources. This method uses a declarative interface between the ontology and the data source, and this interface is modeled in the ontology and implemented using XML schema. Data is imported from the relational source into the ontology using XML, and data integrity is checked by validating the XML submission with an XML schema. We have implemented this approach in PharmGKB (http://www.pharmgkb.org/), a pharmacogenetics knowledge base. Our goals were to (1) import genetic sequence data, collected in relational format, into the pharmacogenetics ontology, and (2) automate the process of updating the links between the ontology and data acquisition when the ontology changes. We tested our approach by linking PharmGKB with data acquisition from a relational model of genetic sequence information. The ontology subsequently evolved, and we were able to rapidly update our interface with the external data and continue acquiring the data. Similar approaches may be helpful for integrating other heterogeneous information sources in order make the diversity of pharmacogenetics data amenable to computational analysis.
PURPOSE: The retinal pigment epithelium (RPE) and choroid comprise a functional unit of the eye that is essential to normal retinal health and function. Here we describe expressed sequence tag (EST) analysis of human RPE/choroid as part of a project for ocular bioinformatics. METHODS: A cDNA library (cs) was made from human RPE/choroid and sequenced. Data were analyzed and assembled using the program GRIST (GRouping and Identification of Sequence Tags). Complete sequencing, Northern and Western blots, RH mapping, peptide antibody synthesis and immunofluorescence (IF) have been used to examine expression patterns and genome location for selected transcripts and proteins. RESULTS: Ten thousand individual sequence reads yield over 6300 unique gene clusters of which almost half have no matches with named genes. One of the most abundant transcripts is from a gene (named "alpha") that maps to the BBS1 region of chromosome 11. A number of tissue preferred transcripts are common to both RPE/choroid and iris. These include oculoglycan/opticin, for which an alternative splice form is detected in RPE/choroid, and "oculospanin" (Ocsp), a novel tetraspanin that maps to chromosome 17q. Antiserum to Ocsp detects expression in RPE, iris, ciliary body, and retinal ganglion cells by IF. A newly identified gene for a zinc-finger protein (TIRC) maps to 19q13.4. Variant transcripts of several genes were also detected. Most notably, the predominant form of Bestrophin represented in cs contains a longer open reading frame as a result of splice junction skipping. CONCLUSIONS: The unamplified cs library gives a view of the transcriptional repertoire of the adult RPE/choroid. A large number of potentially novel genes and splice forms and candidates for genetic diseases are revealed. Clones from this collection are being included in a large, nonredundant set for cDNA microarray construction.
We studied the growth and capacities for pesticides removal of bacterial strains isolated from the Laguna Grande, an oligotrophic lake at the South of Spain (Archidona, Málaga). Strains were isolated from water samples amended with 10 and 50 microg/ml of nine pesticides: organochlorinated insecticides (aldrin and lindane), organophosphorous insecticides (dimetoate, methyl-parathion and methidation), s-triazine herbicides (simazine and atrazine), fungicide (captan) and diflubenzuron (1-(-4-chlorophenyl)-3-(2,6-difluorobenzoyl urea), a chitinase inhibitor. The majority of the strains belonged to the genera Pseudomonas and Aeromonas and only 9% of the total of strains were Gram positive. From all the strains isolated, only 22 showed a wide growth range in all the pesticides tested and 4 of them were chosen for pesticide removal studies. The genetic identification of these strains showed their affiliation to Pseudomonas pseudoalcaligenes, Micrococcus luteus, Bacillus sp. and Exiguobacterium aurantiacum. These last two strains were those that showed the highest pesticide removal capacities and a high bacterial growth.
This article considers how we should frame the ethical issues raised by current proposals for large-scale genebanks with on-going links to medical and lifestyle data, such as the Wellcome Trust and Medical Research Council's 'UK Biobank'. As recent scandals such as Alder Hey have emphasised, there are complex issues concerning the informed consent of donors that need to be carefully considered. However, we believe that a preoccupation with informed consent obscures important questions about the purposes to which such collections are put, not least that they may be only haphazardly used for research (especially that of commercial interest)--an end that would not fairly reflect the original altruistic motivation of donors, and the trust they must invest. We therefore argue that custodians of such databases take on a weighty pro-active duty, to encourage public debate about the ends of such collections and to sponsor research that reflects publicly agreed priorities and provides public benefits.
Recent studies have suggested that a high-density single nucleotide polymorphism (SNP) marker set could provide equivalent or even superior information compared with currently used microsatellite (STR) marker sets for gene mapping by linkage. The focus of this study was to compare results obtained from linkage analyses involving extended pedigrees with STR and single-nucleotide polymorphism (SNP) marker sets. We also wanted to compare the performance of current linkage programs in the presence of high marker density and extended pedigree structures. One replicate of the Genetic Analysis Workshop 14 (GAW14) simulated extended pedigrees (n = 50) from New York City was analyzed to identify the major gene D2. Four marker sets with varying information content and density on chromosome 3 (STR [7.5 cM]; SNP [3 cM, 1 cM, 0.3 cM]) were analyzed to detect two traits, the original affection status, and a redefined trait more closely correlated with D2. Multipoint parametric and nonparametric linkage analyses (NPL) were performed using programs GENEHUNTER, MERLIN, SIMWALK2, and S.A.G.E. SIBPAL. Our results suggested that the densest SNP map (0.3 cM) had the greatest power to detect linkage for the original trait (genetic heterogeneity), with the highest LOD score/NPL score and mapping precision. However, no significant improvement in linkage signals was observed with the densest SNP map compared with STR or SNP-1 cM maps for the redefined affection status (genetic homogeneity), possibly due to the extremely high information contents for all maps. Finally, our results suggested that each linkage program had limitations in handling the large, complex pedigrees as well as a high-density SNP marker set.
OBJECTIVE: The strains of Yersinia pestis isolated in different period and different natural foci in China were analyzed. METHODS: Traditional and molecular biological methods were used. Rhamnose fermentation, rRNA gene copy number, nitrite reduction, and the glycerol fermentation were important characters for typing, and pulse field gel electrophoresis (PFGE) and random amplified polymorphic DNA (RAPD) profile could reflect the genetic distance between the strains. RESULTS: The strains could be divided into 15 genetic types by those 6 characters with each of them covered an isolated geographical territories. CONCLUSION: The characters of strains were described; the genetic relationship of different types, their evolution, and the forming and shift of plague natural foci were analyzed.
Large-scale, multilocus genetic association studies require powerful and appropriate statistical-analysis tools that are designed to relate genotype and haplotype information to phenotypes of interest. Many analysis approaches consider relating allelic, haplotypic, or genotypic information to a trait through use of extensions of traditional analysis techniques, such as contingency-table analysis, regression methods, and analysis-of-variance techniques. In this work, we consider a complementary approach that involves the characterization and measurement of the similarity and dissimilarity of the allelic composition of a set of individuals' diploid genomes at multiple loci in the regions of interest. We describe a regression method that can be used to relate variation in the measure of genomic dissimilarity (or "distance") among a set of individuals to variation in their trait values. Weighting factors associated with functional or evolutionary conservation information of the loci can be used in the assessment of similarity. The proposed method is very flexible and is easily extended to complex multilocus-analysis settings involving covariates. In addition, the proposed method actually encompasses both single-locus and haplotype-phylogeny analysis methods, which are two of the most widely used approaches in genetic association analysis. We showcase the method with data described in the literature. Ultimately, our method is appropriate for high-dimensional genomic data and anticipates an era when cost-effective exhaustive DNA sequence data can be obtained for a large number of individuals, over and above genotype information focused on a few well-chosen loci.
Genomics is the study of the structure and function of the genome: the set of genetic information encoded in the DNA of the nucleus and organelles of an organism. It is a dynamic field that combines traditional paths of inquiry with new approaches that would have been impossible without recent technological developments. Much of the recent focus has been on obtaining the sequence of entire genomes, determining the order and organization of the genes, and developing libraries that provide immediate physical access to any desired DNA fragment. This has enabled functional studies on a genome-wide level, including analysis of the genetic basis of complex traits, quantification of global patterns of gene expression, and systematic gene disruption projects. The successful contribution of genomics to problems in applied entomology requires the cooperation of the private and public sectors to build upon the knowledge derived from the Drosophila genome and effectively develop models for other insect Orders.
OBJECTIVE: To outline the main issues related to the impact of the data generated by the Human Genome Project on health information systems. A major challenge for medical informatics is identified, consisting of adapting traditional systems to new genetic-based diagnostic and therapeutic tools. METHODS: Reviewing and analysing the different health information levels from an organisational complexity point of view. A model is proposed to explain the interactions between health informatics, bioinformatics and molecular medicine. RESULTS: We suggest a new framework that integrates genetic data into health information systems. Using this model, new topics for future research and development are identified. CONCLUSIONS: We are witnessing the birth of a new era (post-genomics). In this era technological advancements in genomics offer new opportunities for clinical applications. Medical informaticians should play an important role in this new endeavour.
Variation in major histocompatibility complex genes on chromosome 6p21.3, specifically the human leukocyte antigen HLA-DR2 or DRB1*1501-DQB1*0602 extended haplotype, confers risk for multiple sclerosis (MS). Previous studies of DRB1 variation and both MS susceptibility and phenotypic expression have lacked statistical power to detect modest genotypic influences, and have demonstrated conflicting results. Results derived from analyses of 1339 MS families indicate DRB1 variation influences MS susceptibility in a complex manner. DRB1*15 was strongly associated in families (P=7.8x10(-31)), and a dominant DRB1*15 dose effect was confirmed (OR=7.5, 95% CI=4.4-13.0, P<0.0001). A modest dose effect was also detected for DRB1*03; however, in contrast to DRB1*15, this risk was recessive (OR=1.8, 95% CI=1.1-2.9, P=0.03). Strong evidence for under-transmission of DRB1*14 (P=5.7x10(-6)) even after accounting for DRB1*15 (P=0.03) was present, confirming a protective effect. In addition, a high risk DRB1*15 genotype bearing DRB1*08 was identified (OR=7.7, 95% CI=4.1-14.4, P<0.0001), providing additional evidence for trans DRB1 allelic interactions in MS. Further, a significant DRB1*15 association observed in primary progressive MS families (P=0.0004), similar to relapsing-remitting MS families, suggests that DRB1-related mechanisms are contributing to both phenotypes. In contrast, results obtained from 2201 MS cases argue convincingly that DRB1*15 genotypes do not modulate age of onset, or significantly influence disease severity measured using expanded disease disability score and disease duration. These results contribute substantially to our understanding of the DRB1 locus and MS, and underscore the importance of using large sample sizes to detect modest genetic effects, particularly in studies of genotype-phenotype relationships.
Because deleterious alleles arising from mutation are filtered by natural selection, mutations that create such alleles will be underrepresented in the set of common genetic variation existing in a population at any given time. Here, we describe an approach based on this idea called VERIFY (variant elimination reinforces functionality), which can be used to assess the extent of natural selection acting on an oligonucleotide motif or set of motifs predicted to have biological activity. As an application of this approach, we analyzed a set of 238 hexanucleotides previously predicted to have exonic splicing enhancer (ESE) activity in human exons using the relative enhancer and silencer classification by unanimous enrichment (RESCUE)-ESE method. Aligning the single nucleotide polymorphisms (SNPs) from the public human SNP database to the chimpanzee genome allowed inference of the direction of the mutations that created present-day SNPs. Analyzing the set of SNPs that overlap RESCUE-ESE hexamers, we conclude that nearly one-fifth of the mutations that disrupt predicted ESEs have been eliminated by natural selection (odds ratio = 0.82 +/- 0.05). This selection is strongest for the predicted ESEs that are located near splice sites. Our results demonstrate a novel approach for quantifying the extent of natural selection acting on candidate functional motifs and also suggest certain features of mutations/SNPs, such as proximity to the splice site and disruption or alteration of predicted ESEs, that should be useful in identifying variants that might cause a biological phenotype.
A single candidate 4'-phosphopantetheine transferase, identified by BLAST searches of the human genome sequence data base, has been cloned, expressed, and characterized. The human enzyme, which is expressed mainly in the cytosolic compartment in a wide range of tissues, is a 329-residue, monomeric protein. The enzyme is capable of transferring the 4'-phosphopantetheine moiety of coenzyme A to a conserved serine residue in both the acyl carrier protein domain of the human cytosolic multifunctional fatty acid synthase and the acyl carrier protein associated independently with human mitochondria. The human 4'-phosphopantetheine transferase is also capable of phosphopantetheinylation of peptidyl carrier and acyl carrier proteins from prokaryotes. The same human protein also has recently been implicated in phosphopantetheinylation of the alpha-aminoadipate semialdehyde dehydrogenase involved in lysine catabolism (Praphanphoj, V., Sacksteder, K. A., Gould, S. J., Thomas, G. H., and Geraghty, M. T. (2001) Mol. Genet. Metab. 72, 336-342). Thus, in contrast to yeast, which utilizes separate 4'-phosphopantetheine transferases to service each of three different carrier protein substrates, humans appear to utilize a single, broad specificity enzyme for all posttranslational 4'-phosphopantetheinylation reactions.