Search PubMed⌕ Search

Biomedical subjects

Toby Johnson

Publications and source records attributed to Toby Johnson.

8 recordsLinked to original sources

Bayesian method for gene detection and mapping, using a case and control design and DNA pooling.

Association mapping studies aim to determine the genetic basis of a trait. A common experimental design uses a sample of unrelated individuals classified into 2 groups, for example cases and controls. If the trait has a complex genetic basis, consisting of many quantitative trait loci (QTLs), each group needs to be large. Each group must be genotyped at marker loci covering the region of interest; for dense coverage of a large candidate region, or a whole-genome scan, the number of markers will be very large. The total amount of genotyping required for such a study is formidable. A laboratory effort efficient technique called DNA pooling could reduce the amount of genotyping required, but the data generated are less informative and require novel methods for efficient analysis. In this paper, a Bayesian statistical analysis of the classic model of McPeek and Strahs is proposed. In contrast to previous work on this model, I assume that data are collected using DNA pooling, so individual genotypes are not directly observed, and also account for experimental errors. A complete analysis can be performed using analytical integration, a propagation algorithm for a hidden Markov model, and quadrature. The method developed here is both statistically and computationally efficient. It allows simultaneous detection and mapping of a QTL, in a large-scale association mapping study, using data from pooled DNA. The method is shown to perform well on data sets simulated under a realistic coalescent-with-recombination model, and is shown to outperform classical single-point methods. The method is illustrated on data consisting of 27 markers in an 880-kb region around the CYP2D6 gene.

Alleles↗

Performance of marker-based relatedness estimators in natural populations of outbred vertebrates.

Knowledge of relatedness between pairs of individuals plays an important role in many research areas including evolutionary biology, quantitative genetics, and conservation. Pairwise relatedness estimation methods based on genetic data from highly variable molecular markers are now used extensively as a substitute for pedigrees. Although the sampling variance of the estimators has been intensively studied for the most common simple genetic relationships, such as unrelated, half- and full-sib, or parent-offspring, little attention has been paid to the average performance of the estimators, by which we mean the performance across all pairs of individuals in a sample. Here we apply two measures to quantify the average performance: first, misclassification rates between pairs of genetic relationships and, second, the proportion of variance explained in the pairwise relatedness estimates by the true population relatedness composition (i.e., the frequencies of different relationships in the population). Using simulated data derived from exceptionally good quality marker and pedigree data from five long-term projects of natural populations, we demonstrate that the average performance depends mainly on the population relatedness composition and may be improved by the marker data quality only within the limits of the population relatedness composition. Our five examples of vertebrate breeding systems suggest that due to the remarkably low variance in relatedness across the population, marker-based estimates may often have low power to address research questions of interest.

Animals↗

MCALIGN2: faster, accurate global pairwise alignment of non-coding DNA sequences based on explicit models of indel evolution.

BACKGROUND: Non-coding DNA sequences comprise a very large proportion of the total genomic content of mammals, most other vertebrates, many invertebrates, and most plants. Unraveling the functional significance of non-coding DNA depends on how well we are able to align non-coding DNA sequences. However, the alignment of non-coding DNA sequences is more difficult than aligning protein-coding sequences. RESULTS: Here we present an improved pair-hidden-Markov-Model (pair HMM) based method for performing global pairwise alignment of non-coding DNA sequences. The method uses an explicit model of indel length frequency distribution which can be specified, and allows any time reversible model of nucleotide substitution. The method uses a deterministic global optimiser to find the alignment with the highest posterior probability. We test MCALIGN2 in simulations, and compare it to a previous Monte Carlo based method (MCALIGN), to the pair HMM method of Knudsen and Miyamoto, and to a heuristic method (AVID) that performed very well in a previous simulation study. We show that the pair HMM methods have excellent performance for all combinations of parameter values we have considered. MCALIGN2 is up to ten times faster than MCALIGN. MCALIGN2 is more accurate in resolving indels given an accurate explicit model than heuristic methods, but is computationally slower. CONCLUSION: MCALIGN2 produces better quality alignments by explicitly using biological knowledge about the indel length distribution and time reversible models of nucleotide substitution. As a result, it can outperform other available sequence alignment methods for the cases we have considered to align non-coding DNA sequences.

Algorithms↗

Theoretical models of selection and mutation on quantitative traits.

Empirical studies of quantitative genetic variation have revealed robust patterns that are observed both across traits and across species. However, these patterns have no compelling explanation, and some of the observations even appear to be mutually incompatible. We review and extend a major class of theoretical models, 'mutation-selection models', that have been proposed to explain quantitative genetic variation. We also briefly review an alternative class of 'balancing selection models'. We consider to what extent the models are compatible with the general observations, and argue that a key issue is understanding and modelling pleiotropy. We discuss some of the thorny issues that arise when formulating models that describe many traits simultaneously.

Evolution, Molecular↗

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms↗

The fixation probability of a beneficial allele in a population dividing by binary fission.

We derive formulae for the fixation probability, P, of a rare benefical allele segregating in a population of fixed size which reproduces by binary fission, in terms of the selection coefficient for the beneficial allele, s. We find that an earlier result P approximately = 4s does not depend on the assumption of binary fission, but depends on an assumption about the ordering of events in the life cycle. We find that P approximately = 2s for mutations occurring during chromosome replication and P approximately = 2.8s for mutations occurring at random times between replication events.

Alleles↗

General models of multilocus evolution.

In 1991, Barton and Turelli developed recursions to describe the evolution of multilocus systems under arbitrary forms of selection. This article generalizes their approach to allow for arbitrary modes of inheritance, including diploidy, polyploidy, sex linkage, cytoplasmic inheritance, and genomic imprinting. The framework is also extended to allow for other deterministic evolutionary forces, including migration and mutation. Exact recursions that fully describe the state of the population are presented; these are implemented in a computer algebra package (available on the Web at http://helios.bto.ed.ac.uk/evolgen). Despite the generality of our framework, it can describe evolutionary dynamics exactly by just two equations. These recursions can be further simplified using a "quasi-linkage equilibrium" (QLE) approximation. We illustrate the methods by finding the effect of natural selection, sexual selection, mutation, and migration on the genetic composition of a population.

Emigration and Immigration↗

The effect of deleterious alleles on adaptation in asexual populations.

We calculate the fixation probability of a beneficial allele that arises as the result of a unique mutation in an asexual population that is subject to recurrent deleterious mutation at rate U. Our analysis is an extension of previous works, which make a biologically restrictive assumption that selection against deleterious alleles is stronger than that on the beneficial allele of interest. We show that when selection against deleterious alleles is weak, beneficial alleles that confer a selective advantage that is small relative to U have greatly reduced probabilities of fixation. We discuss the consequences of this effect for the distribution of effects of alleles fixed during adaptation. We show that a selective sweep will increase the fixation probabilities of other beneficial mutations arising during some short interval afterward. We use the calculated fixation probabilities to estimate the expected rate of fitness improvement in an asexual population when beneficial alleles arise continually at some low rate proportional to U. We estimate the rate of mutation that is optimal in the sense that it maximizes this rate of fitness improvement. Again, this analysis relaxes the assumption made previously that selection against deleterious alleles is stronger than on beneficial alleles.

Adaptation, Physiological↗