Search PubMed⌕ Search

Biomedical subjects

D J Balding

Publications and source records attributed to D J Balding.

At least 19 recordsLinked to original sources

Fine-scale mapping of disease loci via shattered coalescent modeling of genealogies.

We present a Bayesian, Markov-chain Monte Carlo method for fine-scale linkage-disequilibrium gene mapping using high-density marker maps. The method explicitly models the genealogy underlying a sample of case chromosomes in the vicinity of a putative disease locus, in contrast with the assumption of a star-shaped tree made by many existing multipoint methods. Within this modeling framework, we can allow for missing marker information and for uncertainty about the true underlying genealogy and the makeup of ancestral marker haplotypes. A crucial advantage of our method is the incorporation of the shattered coalescent model for genealogies, allowing for multiple founding mutations at the disease locus and for sporadic cases of disease. Output from the method includes approximate posterior distributions of the location of the disease locus and population-marker haplotype proportions. In addition, output from the algorithm is used to construct a cladogram to represent genetic heterogeneity at the disease locus, highlighting clusters of case chromosomes sharing the same mutation. We present detailed simulations to provide evidence of improvements over existing methodology. Furthermore, inferences about the location of the disease locus are shown to remain robust to modeling assumptions.

Algorithms↗

Measuring gametic disequilibrium from multilocus data.

We describe a Bayesian approach to analyzing multilocus genotype or haplotype data to assess departures from gametic (linkage) equilibrium. Our approach employs a Markov chain Monte Carlo (MCMC) algorithm to approximate the posterior probability distributions of disequilibrium parameters. The distributions are computed exactly in some simple settings. Among other advantages, posterior distributions can be presented visually, which allows the uncertainties in parameter estimates to be readily assessed. In addition, background knowledge can be incorporated, where available, to improve the precision of inferences. The method is illustrated by application to previously published datasets; implications for multilocus forensic match probabilities and for simple association-based gene mapping are also discussed.

Algorithms↗

Models of sequence evolution for DNA sequences containing gaps.

Most evolutionary tree estimation methods for DNA sequences ignore or inefficiently use the phylogenetic information contained within shared patterns of gaps. This is largely due to the computational difficulties in implementing models for insertions and deletions. A simple way to incorporate this information is to treat a gap as a fifth character (with the four nucleotides being the other four) and to incorporate it within a Markov model of nucleotide substitution. This idea has been dismissed in the past, since it treats a multiple-site insertion or deletion as a sequence of independent events rather than a single event. While this is true, we have found that under many circumstances it is better to incorporate gap information inadequately than to ignore it, at least for topology estimation. We propose an extension to a class of nucleotide substitution models to incorporate the gap character and show that, for data sets (both real and simulated) with short and medium gaps, these models do lead to effective use of the information contained within insertions and deletions. We also implement an ad hoc method in which the likelihood at columns containing multiple-site gaps is downweighted in order to avoid giving them undue influence. The precision of the estimated tree, assessed using Markov chain Monte Carlo techniques to find the posterior distribution over tree space, improves under these five-state models compared with standard methods which effectively ignore gaps.

Algorithms↗

Bayesian fine-scale mapping of disease loci, by hidden Markov models.

We present a new multilocus method for the fine-scale mapping of genes contributing to human diseases. The method is designed for use with multiple biallelic markers-in particular, single-nucleotide polymorphisms for which high-density genetic maps will soon be available. We model disease-marker association in a candidate region via a hidden Markov process and allow for correlation between linked marker loci. Using Markov-chain-Monte Carlo simulation methods, we obtain posterior distributions of model parameter estimates including disease-gene location and the age of the disease-predisposing mutation. In addition, we allow for heterogeneity in recombination rates, across the candidate region, to account for recombination hot and cold spots. We also obtain, for the ancestral marker haplotype, a posterior distribution that is unique to our method and that, unlike maximum-likelihood estimation, can properly account for uncertainty. We apply the method to data for cystic fibrosis and Huntington disease, for which mutations in disease genes have already been identified. The new method performs well compared with existing multi-locus mapping methods.

Alleles↗

Measuring departures from Hardy-Weinberg: a Markov chain Monte Carlo method for estimating the inbreeding coefficient.

Many well-established statistical methods in genetics were developed in a climate of severe constraints on computational power. Recent advances in simulation methodology now bring modern, flexible statistical methods within the reach of scientists having access to a desktop workstation. We illustrate the potential advantages now available by considering the problem of assessing departures from Hardy-Weinberg (HW) equilibrium. Several hypothesis tests of HW have been established, as well as a variety of point estimation methods for the parameter which measures departures from HW under the inbreeding model. We propose a computational, Bayesian method for assessing departures from HW, which has a number of important advantages over existing approaches. The method incorporates the effects-of uncertainty about the nuisance parameters--the allele frequencies--as well as the boundary constraints on f (which are functions of the nuisance parameters). Results are naturally presented visually, exploiting the graphics capabilities of modern computer environments to allow straightforward interpretation. Perhaps most importantly, the method is founded on a flexible, likelihood-based modelling framework, which can incorporate the inbreeding model if appropriate, but also allows the assumptions of the model to he investigated and, if necessary, relaxed. Under appropriate conditions, information can be shared across loci and, possibly, across populations, leading to more precise estimation. The advantages of the method are illustrated by application both to simulated data and to data analysed by alternative methods in the recent literature.

Algorithms↗

Genealogical inference from microsatellite data.

Ease and accuracy of typing, together with high levels of polymorphism and widespread distribution in the genome, make microsatellite (or short tandem repeat) loci an attractive potential source of information about both population histories and evolutionary processes. However, microsatellite data are difficult to interpret, in particular because of the frequency of back-mutations. Stochastic models for the underlying genetic processes can be specified, but in the past they have been too complicated for direct analysis. Recent developments in stochastic simulation methodology now allow direct inference about both historical events, such as genealogical coalescence times, and evolutionary parameters, such as mutation rates. A feature of the Markov chain Monte Carlo (MCMC) algorithm that we propose here is that the likelihood computations are simplified by treating the (unknown) ancestral allelic states as auxiliary parameters. We illustrate the algorithm by analyzing microsatellite samples simulated under the model. Our results suggest that a single microsatellite usually does not provide enough information for useful inferences, but that several completely linked microsatellites can be informative about some aspects of genealogical history and evolutionary processes. We also reanalyze data from a previously published human Y chromosome microsatellite study, finding evidence for an effective population size for human Y chromosomes in the low thousands and a recent time since their most recent common ancestor: the 95% interval runs from approximately 15, 000 to 130,000 years, with most likely values around 30,000 years.

Genetics, Population↗

The design of pooling experiments for screening a clone map.

We consider nonadaptive pooling designs for unique-sequence screening of a 1530-clone map of Aspergillus nidulans. The map has the properties that the clones are, with possibly a few exceptions, ordered and no more than 2 of them cover any point on the genome. We propose two subdesigns of the Steiner system S(3, 5, 65), one with 65 pools and approximately 118 clones per pool, the other with 54 pools and about 142 clones per pool. Each design allows 1 or 2 positive clones to be detected, even in the presence of substantial experimental error rates. More efficient designs are possible if the overlap information in the map is exploited, if there is no constraint on the number of clones in a pool, and if no error tolerance is required. An information theory lower bound requires at least 12 pools to satisfy these minimal criteria, and an "interleaved binary" design can be constructed on 20 pools, with about 380 clones per pool. However, the designs with more pools have important properties of robustness to various possible errors and general applicability to a wider class of pooling experiments.

Aspergillus nidulans↗

Significant genetic correlations among Caucasians at forensic DNA loci.

Although the effect of population differentiation on the forensic use of DNA profiles has been the subject of controversy for some years now, the debate has largely failed to focus on the genetical questions directly relevant to the forensic context. We re-analyse two published data sets and find that they convey much the same message for forensic inference, in contrast with the dramatically differing conclusions of the original authors. The analysis is likelihood-based and combines information across loci and across populations without assuming constant genetic differentiation. Our results suggest that the relevant genetic correlation coefficients are too large to be ignored in forensic work: although DNA profile evidence is typically very strong, the effect of genetic correlations can be important in some cases. Such correlations can, however, be accommodated in an appropriate assessment of evidential strength so that population genetic issues should not present a barrier to the efficient and fair use of DNA profile evidence.

DNA Fingerprinting↗

Inferring coalescence times from DNA sequence data.

The paper is concerned with methods for the estimation of the coalescence time (time since the most recent common ancestor) of a sample of intraspecies DNA sequences. The methods take advantage of prior knowledge of population demography, in addition to the molecular data. While some theoretical results are presented, a central focus is on computational methods. These methods are easy to implement, and, since explicit formulae tend to be either unavailable or unilluminating, they are also more useful and more informative in most applications. Extensions are presented that allow for the effects of uncertainty in our knowledge of population size and mutation rates, for variability in population sizes, for regions of different mutation rate, and for inference concerning the coalescence time of the entire population. The methods are illustrated using recent data from the human Y chromosome.

Algorithms↗

Population genetics of STR loci in Caucasians.

STR loci are becoming increasingly important in forensic casework. In order to be used fairly and efficiently, the population genetics of these loci must be investigated and the implications for forensic inference assessed. A key population genetics parameter is the "coancestry coefficient", or FST, which is the correlation between two genes sampled from distinct individuals within a subpopulation. We present analyses of STR data, at geographic scales which range from national to regional, from the UK and other European sources. We implement a likelihood-based method of estimating FST, which has important advantages over alternative methods: it allows a range of plausible values to be assessed, rather than presenting a single point estimate, and it allows a subpopulation to be compared with a larger population from which a database has been drawn, which is the relevant comparison in forensic work. Our results suggest that values of FST appropriate to forensic applications in Europe are too large to be ignored. With appropriate allowance, however, it is possible to make use of STR evidence in a way which is efficient yet avoids overstatement of evidential strength.

Cross-Cultural Comparison↗

Evaluating DNA profile evidence when the suspect is identified through a database search.

The paper is concerned with the strength of DNA evidence when a suspect is identified via a search through a database of the DNA profiles of known individuals. Consideration of the appropriate likelihood ratio shows that in this setting the DNA evidence is (slightly) stronger than when a suspect is identified by other means, subsequently profiled, and found to match. The recommendation of the 1992 report of the US National Research Council that DNA evidence that is used to identify the suspect should not be presented at trial thus seems unnecessarily conservative. The widely held view that DNA evidence is weaker when it results from a database search seems to be based on a rationale that leads to absurd conclusions in some examples. Moreover, this view is inconsistent with the principle, which enjoys substantial support, that evidential weight should be measured by likelihood ratios. The strength of DNA evidence is shown also to be slightly increased for other forms of search procedure. While the DNA evidence is stronger after a database search, the overall case against the suspect may not be, and the problems of incorporating the DNA with the non-DNA evidence can be particularly important in such cases.

Crime↗

Inferring identify from DNA profile evidence.

The controversy over the interpretation of DNA profile evidence in forensic identification can be attributed in part to confusion over the mode(s) of statistical inference appropriate to this setting. Although there has been substantial discussion in the literature of, for example, the role of population genetics issues, few authors have made explicit the inferential framework which underpins their arguments. This lack of clarity has led both to unnecessary debates over ill-posed or inappropriate questions and to the neglect of some issues which can have important consequences. We argue that the mode of statistical inference which seems to underlie the arguments of some authors, based on a hypothesis testing framework, is not appropriate for forensic identification. We propose instead a logically coherent framework in which, for example, the roles both of the population genetics issues and of the nonscientific evidence in a case are incorporated. Our analysis highlights several widely held misconceptions in the DNA profiling debate. For example, the profile frequency is not directly relevant to forensic inference. Further, very small match probabilities may in some settings be consistent with acquittal. Although DNA evidence is typically very strong, our analysis of the coherent approach highlights situations which can arise in practice where alternative methods for assessing DNA evidence may be misleading.

Criminal Law↗

Efficient pooling designs for library screening.

We describe efficient methods for screening clone libraries, based on pooling schemes that we call "random k-sets designs." In these designs, the pools in which any clone occurs are equally likely to be any possible selection of k from the v pools. The values of k and v can be chosen to optimize desirable properties. Random k-sets designs have substantial advantages over alternative pooling schemes: they are efficient, flexible, and easy to specify, require fewer pools, and have error-correcting and error-detecting capabilities. In addition, screening can often be achieved in only one pass, thus facilitating automation. For design comparison, we assume a binomial distribution for the number of "positive" clones, with parameters n, the number of clones, and c, the coverage. We propose the expected number of resolved positive clones--clones that are definitely positive based upon the pool assays--as a criterion for the efficiency of a pooling design. We determine the value of k that is optimal, with respect to this criterion, as a function of v, n, and c. We also describe superior k-sets designs called k-sets packing designs. As an illustration, we discuss a robotically implemented design for a 2.5-fold-coverage, human chromosome 16 YAC library of n = 1298 clones. We also estimate the probability that each clone is positive, given the pool-assay data and a model for experimental errors.

Binomial Distribution↗

A method for quantifying differentiation between populations at multi-allelic loci and its implications for investigating identity and paternity.

A method is proposed for allowing for the effects of population differentiation, and other factors, in forensic inference based on DNA profiles. Much current forensic practice ignores, for example, the effects of coancestry and inappropriate databases and is consequently systematically biased against defendants. Problems with the 'product rule' for forensic identification have been highlighted by several authors, but important aspects of the problems are not widely appreciated. This arises in part because the match probability has often been confused with the relative frequency of the profile. Further, the analogous problems in paternity cases have received little attention. The proposed method is derived under general assumptions about the underlying population genetic processes. Probabilities relevant to forensic inference are expressed in terms of a single parameter whose values can be chosen to reflect the specific circumstances. The method is currently used in some UK courts and has important advantages over the 'Ceiling Principle' method, which has been criticized on a number of grounds.

Alleles↗

Design and analysis of chromosome physical mapping experiments.

Mathematical and statistical aspects of constructing ordered-clone physical maps of chromosomes are reviewed. Three broad problems are addressed: analysis of fingerprint data to identify configurations of overlapping clones, prediction of the rate of progress of a mapping strategy and optimal design of pooling schemes for screening large clone libraries.

Chromosome Mapping↗