Search PubMed⌕ Search

Biomedical subjects

John Molitor

Publications and source records attributed to John Molitor.

12 recordsLinked to original sources

Bayesian modeling of air pollution health effects with missing exposure data.

The authors propose a new statistical procedure that utilizes measurement error models to estimate missing exposure data in health effects assessment. The method detailed in this paper follows a Bayesian framework that allows estimation of various parameters of the model in the presence of missing covariates in an informative way. The authors apply this methodology to study the effect of household-level long-term air pollution exposures on lung function for subjects from the Southern California Children's Health Study pilot project, conducted in the year 2000. Specifically, they propose techniques to examine the long-term effects of nitrogen dioxide (NO2) exposure on children's lung function for persons living in 11 southern California communities. The effect of nitrogen dioxide exposure on various measures of lung function was examined, but, similar to many air pollution studies, no completely accurate measure of household-level long-term nitrogen dioxide exposure was available. Rather, community-level nitrogen dioxide was measured continuously over many years, but household-level nitrogen dioxide exposure was measured only during two 2-week periods, one period in the summer and one period in the winter. From these incomplete measures, long-term nitrogen dioxide exposure and its effect on health must be inferred. Results show that the method improves estimates when compared with standard frequentist approaches.

Air Pollutants↗

Association mapping with single-feature polymorphisms.

We develop methods for exploiting "single-feature polymorphism" data, generated by hybridizing genomic DNA to oligonucleotide expression arrays. Our methods enable the use of such data, which can be regarded as very high density, but imperfect, polymorphism data, for genomewide association or linkage disequilibrium mapping. We use a simulation-based power study to conclude that our methods should have good power for organisms like Arabidopsis thaliana, in which linkage disequilibrium is extensive, the reason being that the noisiness of single-feature polymorphism data is more than compensated for by their great number. Finally, we show how power depends on the accuracy with which single-feature polymorphisms are called.

Algorithms↗

Fine mapping--19th century style.

BACKGROUND: There is great interest in the use of computationally intensive methods for fine mapping of marker data. In this paper we develop methods based upon ideas originally proposed 100 years ago in the context of spatial clustering. METHODS: We use spatial clustering of haplotypes as a low-dimensional surrogate for the unobserved genealogy underlying a set of genotype data. In doing so we hope to avoid the computational complexity inherent in explicitly modelling details of the ancestry of the sample, while at the same time capturing the key correlations induced by that ancestry at a much lower computational cost. RESULTS: We benchmark our methods using the simulated Genetic Analysis Workshop 14 data, using 100 replicates of 4 phenotypes to indicate the power of our method. When a functional mutation relating to a trait is actually present, we find evidence for that mutation in 97 out of 100 replicates, on average. CONCLUSION: Our results show that our method has the ability to accurately infer the location of functional mutations from unphased genotype data.

Congresses as Topic↗

Genome-wide association mapping in Arabidopsis identifies previously known flowering time and pathogen resistance genes.

There is currently tremendous interest in the possibility of using genome-wide association mapping to identify genes responsible for natural variation, particularly for human disease susceptibility. The model plant Arabidopsis thaliana is in many ways an ideal candidate for such studies, because it is a highly selfing hermaphrodite. As a result, the species largely exists as a collection of naturally occurring inbred lines, or accessions, which can be genotyped once and phenotyped repeatedly. Furthermore, linkage disequilibrium in such a species will be much more extensive than in a comparable outcrossing species. We tested the feasibility of genome-wide association mapping in A. thaliana by searching for associations with flowering time and pathogen resistance in a sample of 95 accessions for which genome-wide polymorphism data were available. In spite of an extremely high rate of false positives due to population structure, we were able to identify known major genes for all phenotypes tested, thus demonstrating the potential of genome-wide association mapping in A. thaliana and other species with similar patterns of variation. The rate of false positives differed strongly between traits, with more clinal traits showing the highest rate. However, the false positive rates were always substantial regardless of the trait, highlighting the necessity of an appropriate genomic control in association studies.

Arabidopsis↗

Effects of simultaneous over-expression of Cu/ZnSOD and MnSOD on Drosophila melanogaster life span.

The FLP-out technique, based on yeast FLP recombinase, allows induced over-expression of transgenes in Drosophila adults. With FLP-out control and over-expressing flies have identical genetic backgrounds and therefore differences in life span must result from transgene induction. The amount of over-expression achieved varies between independent transgenic lines, and previously for both Cu/ZnSOD and MnSOD life span was found to be increased in proportion to the increase in enzyme activity. To determine if greater increases in enzyme and life span could be achieved with FLP-out, enzyme over-expression and life span were analyzed in eight lines containing two MnSOD transgenes, three lines containing three MnSOD transgenes, and three lines containing a MnSOD transgene plus a Cu/ZnSOD transgene. Life span was again found to be increased in proportion to the increase in MnSOD enzyme activity, with increases of up to 40% in mean and maximum life span. However the increases in enzyme activity and life span conferred per transgene were reduced when more than one transgene was present at the same time. When the reduced efficiency of enzyme over-expression per transgene was taken into account, simultaneous over-expression of MnSOD and Cu/ZnSOD was found to have partially additive effects on life span.

Animals↗

A survey of current Bayesian gene mapping methods.

Recently, there has been much interest in the use of Bayesian statistical methods for performing genetic analyses. Many of the computational difficulties previously associated with Bayesian analysis, such as multidimensional integration, can now be easily overcome using modern high-speed computers and Markov chain Monte Carlo (MCMC) methods. Much of this new technology has been used to perform gene mapping, especially through the use of multi-locus linkage disequilibrium techniques. This review attempts to summarise some of the currently available methods and the software available to implement these methods.

Bayes Theorem↗

Haplotype structure and phenotypic associations in the chromosomal regions surrounding two Arabidopsis thaliana flowering time loci.

The feasibility of using linkage disequilbrium (LD) to fine-map loci underlying natural variation in Arabidopsis thaliana was investigated by looking for associations between flowering time and marker polymorphism in the genomic regions containing two candidate genes, FRI and FLC, both of which are known to contribute to natural variation in flowering. A sample of 196 accessions was used, and polymorphism was assessed by sequencing a total of 17 roughly 500-bp fragments. Using a novel Bayesian algorithm based on haplotype similarity, we demonstrate that LD could have been used to fine-map the FRI gene to a roughly 30-kb region and to identify two common loss-of-function alleles. Interestingly, because of genetic heterogeneity, simple single-marker associations would not have been able to map FRI with nearly the same precision. No clear evidence for previously unknown alleles at either locus was found, but the effect of population structure in causing false positives was evident.

Arabidopsis↗

Markov chain Monte Carlo without likelihoods.

Many stochastic simulation approaches for generating observations from a posterior distribution depend on knowing a likelihood function. However, for many complex probability models, such likelihoods are either impossible or computationally prohibitive to obtain. Here we present a Markov chain Monte Carlo method for generating observations from a posterior distribution without the use of likelihoods. It can also be used in frequentist applications, in particular for maximum-likelihood estimation. The approach is illustrated by an example of ancestral inference in population genetics. A number of open problems are highlighted in the discussion.

Algorithms↗

Fine-scale mapping of disease genes with multiple mutations via spatial clustering techniques.

We present a method to perform fine mapping by placing haplotypes into clusters on the basis of risk. Each cluster has a haplotype "center." Cluster allocation is defined according to haplotype centers, with each haplotype assigned to the cluster with the "closest" center. The closeness of two haplotypes is determined by a similarity metric that measures the length of the shared segment around the location of a putative functional mutation for the particular cluster. Our method allows for missing marker information but still estimates the risks of complete haplotypes without resorting to a one-marker-at-a-time analysis. The dimensionality issues that can occur in haplotype analyses are removed by sampling over the haplotype space, allowing for estimation of haplotype risks without explicitly assigning a parameter to each haplotype to be estimated. In this way, we are able to handle haplotypes of arbitrary size. Furthermore, our clustering approach has the potential to allow us to detect the presence of multiple functional mutations.

Algorithms↗

Application of Bayesian spatial statistical methods to analysis of haplotypes effects and gene mapping.

We propose a method to analyze haplotype effects using ideas derived from Bayesian spatial statistics. We assume that two haplotypes that are similar to one another in structure are likely to have similar risks, and define a distance metric to specify the appropriate level of closeness between the two haplotypes. Through the choice of distance metric, varying levels of population genetics theory can be incorporated into the modeling process, including some that allow estimation of the location of the disease causing mutation(s). This location can be estimated, along with the other parameters of the model, using Markov chain Monte Carlo (MCMC) estimation methods. We demonstrate the effectiveness of the model on two real datasets, a well-known dataset used to fine-map the gene for cystic fibrosis, and one used to localize the gene for Friedreich's ataxia.

Bayes Theorem↗

Bayesian spatial modeling of haplotype associations.

We review methods for relating the risk of disease to a collection of single nucleotide polymorphisms (SNPs) within a small region. Association studies using case-control designs with unrelated individuals could be used either to test for a direct effect of a candidate gene and characterize the responsible variant(s), or to fine map an unknown gene by exploiting the pattern of linkage disequilibrium (LD). We consider a flexible class of logistic penetrance models based on haplotypes and compare them with an alternative formulation based on unphased multilocus genotypes. The likelihood for haplotype-based models requires summation over all possible haplotype assignments consistent with the observed genotype data, and can be fitted using either Expectation-Maximization (E-M) or Markov chain Monte Carlo (MCMC) methods. Subtleties involving ascertainment correction for case-control studies are discussed. There has been great interest in methods for LD mapping based on the coalescent or ancestral recombination graphs as well as methods based on haplotype sharing, both of which we review briefly. Because of their computational complexity, we propose some alternative empirical modeling approaches using techniques borrowed from the Bayesian spatial statistics literature. Here, space is interpreted in terms of a distance metric describing the similarity of any pair of haplotypes to each other, and hence their presumed common ancestry. Specifically, we discuss the conditional autoregressive model and two spatial clustering models: Potts and Voronoi. We conclude with a discussion of the implications of these methods for modeling cryptic relatedness, haplotype blocks, and haplotype tagging SNPs, and suggest a Bayesian framework for the HapMap project.

Algorithms↗

Bayesian modeling of complex metabolic pathways.

Many chronic diseases are the result of a complex sequence of biochemical reactions involving exposures to various environmental agents, metabolized by a number of different genes. Routine epidemiologic analyses of such associations have tended to rely on standard contingency table or logistic regression methods, typically focusing on one variable at a time or pairwise combinations. We consider two statistical alternatives to this approach, one based on Bayesian model averaging, one based on pharmacokinetic modeling of the biochemical pathways. These approaches are illustrated using data from a case-control study of colorectal polyps in relation to tobacco smoking and consumption of well done red meat, both viewed as sources of heterocyclic amines and polycyclic aromatic hydrocarbons. The new analyses are structured in a manner that attempts to take advantage of prior knowledge of the metabolism of these classes of compounds and the various genes that regulate these pathways.

Bayes Theorem↗