Search PubMed⌕ Search

Biomedical subjects

David J Balding

Publications and source records attributed to David J Balding.

14 recordsLinked to original sources

Family-based association analysis with ordered categorical phenotypes, covariates and interactions.

Genetic association analyses of family-based studies with ordered categorical phenotypes are often conducted using methods either for quantitative or for binary traits, which can lead to suboptimal analyses. Here we present an alternative likelihood-based method of analysis for single nucleotide polymorphism (SNP) genotypes and ordered categorical phenotypes in nuclear families of any size. Our approach, which extends our previous work for binary phenotypes, permits straightforward inclusion of covariate, gene-gene and gene-covariate interaction terms in the likelihood, incorporates a simple model for ascertainment and allows for family-specific effects in the hypothesis test. Additionally, our method produces interpretable parameter estimates and valid confidence intervals. We assess the proposed method using simulated data, and apply it to a polymorphism in the c-reactive protein (CRP) gene typed in families collected to investigate human systemic lupus erythematosus. By including sex interactions in the analysis, we show that the polymorphism is associated with anti-nuclear autoantibody (ANA) production in females, while there appears to be no effect in males.

Autoantibodies↗

Clinical factors and ABCB1 polymorphisms in prediction of antiepileptic drug response: a prospective cohort study.

BACKGROUND: The ABCB1 3435C-->T single-nucleotide polymorphism (SNP) or a three-SNP haplotype containing 3435C-->T has been implicated in multidrug resistance in epilepsy in three retrospective case-control studies, but a further three have failed to replicate the association. We aimed to determine the effect of the ABCB1 gene on epilepsy drug response, using a unique large cohort of epilepsy patients with prospectively measured seizure and drug response outcomes. METHODS: The ABCB1 3435C-->T polymorphism and three-SNP haplotype, plus a comprehensive set of tag SNPs across ABCB1 and adjacent ABCB4, were genotyped in a cohort of 503 epilepsy patients with prospectively measured seizure and drug response outcomes. Clinical, demographic, and genetic data were analysed. Treatment outcome was measured in terms of time to 12-month remission, time to first seizure, and time to drug withdrawal due to inadequate seizure control or side-effects. Randomly selected genome-wide HapMap SNPs (n=129) were genotyped in all patients for genomic control. FINDINGS: Number of seizures before treatment was the dominant feature predicting seizure outcome after starting antiepileptic drug therapy, measured by both time to first seizure (hazard ratio 1.34, 95% CI 1.21-1.49, p<0.0001) and time to 12-month remission (0.83, 0.73-0.94, p=0.003). There was no association of the ABCB1 3435C-->T polymorphism, the three-SNP haplotype, or any gene-wide tag SNP with time to first seizure after starting drug therapy, time to 12-month remission, or time to drug withdrawal due to unacceptable side-effects or to lack of seizure control. INTERPRETATION: We found no evidence that ABCB1 common variation influences either seizure or drug withdrawal outcomes after initiation of antiepileptic drug therapy.

ATP-Binding Cassette Transporters↗

A tutorial on statistical methods for population association studies.

Although genetic association studies have been with us for many years, even for the simplest analyses there is little consensus on the most appropriate statistical procedures. Here I give an overview of statistical approaches to population association studies, including preliminary analyses (Hardy-Weinberg equilibrium testing, inference of phase and missing data, and SNP tagging), and single-SNP and multipoint tests for association. My goal is to outline the key methods with a brief discussion of problems (population structure and multiple testing), avenues for solutions and some ongoing developments.

Genetics, Medical↗

Exon sequencing and high resolution haplotype analysis of ABC transporter genes implicated in drug resistance.

BACKGROUND: The ATP-binding cassette (ABC) proteins are a superfamily of efflux pumps implicated as a mechanism for multidrug resistance in cytotoxic chemotherapy, immunosuppressive therapy, HIV and epilepsy. Genetic variation in P-glycoprotein, the product of the ABCB1 gene, is proposed to mediate de novo drug resistance, but associations between polymorphisms in ABCB1 and pharmacoresistance have produced conflicting results. Potential explanations for the inconsistency of results include inadequate characterization of gene structure, variation and linkage disequilibrium (LD) in ABCB1, as well as overlap in substrate specificity between ABCB1 and the various other drug transporters. METHODS AND RESULTS: We undertook a fundamental analysis of gene structure, variation and LD in ABCB1 and four other drug transporter genes implicated in pharmacoresistance: ABCC1, ABCC2, ABCC5 and ABCB4. Manual annotation of the five genes revealed nine shorter alternative transcripts with new untranslated regions and one novel region of coding sequence, demonstrating that on-line annotations are incomplete. Sequencing of exons in 47 Caucasian individuals identified 75 novel single nucleotide polymorphisms (SNPs) previously undescribed in any public database, including 14 new coding sequence SNPs. Genotyping of 502 SNPs in 842 Caucasian individuals across the five genes revealed large blocks of high LD, and low haplotype diversity across all five genes that could be characterized by between 67 and 114 tagging SNPs, depending on the tagging criteria. CONCLUSION: The study illustrates that publicly available data resources on genomic organization of genes and common variation can have important gaps and limitations, and establishes a comprehensive set of tagging SNPs for future association studies in pharmacoresistance.

ATP-Binding Cassette Transporters↗

Logistic regression protects against population structure in genetic association studies.

We conduct an extensive simulation study to compare the merits of several methods for using null (unlinked) markers to protect against false positives due to cryptic substructure in population-based genetic association studies. The more sophisticated "structured association" methods perform well but are computationally demanding and rely on estimating the correct number of subpopulations. The simple and fast "genomic control" approach can lose power in certain scenarios. We find that procedures based on logistic regression that are flexible, computationally fast, and easy to implement also provide good protection against the effects of cryptic substructure, even though they do not explicitly model the population structure.

Animals↗

Discrimination of half-siblings when maternal genotypes are known.

Given the DNA profiles of two individuals and one parent (say the mother) of each, we present likelihood ratios (LRs) comparing the hypothesis that they have the same father with the hypothesis of unrelated fathers. If the individuals have the same mother, the problem is to distinguish full- from half-siblings, otherwise we are comparing a half-sibling relationship with unrelated. We simulate STR profiles at up to 60 loci, based on allele proportions observed at 15 loci in three populations, and use them to approximate misclassification rates both for binary classification (e.g. "half-sib" versus "unrelated"), and when a third "cannot say" category is included. We find that reliable inferences in the absence of the mothers' profiles require many more STR loci than the 10-25 loci that are currently routinely available. However, profiling the two mothers conveys more discriminatory power than profiling the same number of additional loci in the individuals themselves. Our likelihood ratio formulas include a theta (or Fst) adjustment to allow for the individuals concerned to have recent shared ancestry (coancestry), relative to the population from which the allele frequency database is drawn. We illustrate that using an appropriate value of theta can reduce the average misclassification rate.

DNA Fingerprinting↗

Clustering of protein domains in the human genome.

We present a systematic study of the clustering of genes within the human genome based on homology inferred from both sequence and structural similarity. The 3D-Genomics automated proteome annotation pipeline () was utilised to infer homology for each protein domain in the genome, for the 26 superfamilies most highly represented in the Structural Classification Of Proteins (SCOP) database. This approach enabled us to identify homologues that could not be detected by sequence-based methods alone. For each superfamily, we investigated the distribution, both within and among chromosomes, of genes encoding at least one domain within the superfamily. The results indicate a diversity of clustering behaviours: some superfamilies showed no evidence of any clustering, and others displayed significant clustering either within or among chromosomes, or both. Removal of tandem repeats reduced the levels of clustering observed, but some superfamilies still displayed highly significant clustering. Thus, our study suggests that either the process of gene duplication, or the evolution of the resulting clusters, differs between structural superfamilies.

Cadherins↗

Identifying adaptive genetic divergence among populations from genome scans.

The identification of signatures of natural selection in genomic surveys has become an area of intense research, stimulated by the increasing ease with which genetic markers can be typed. Loci identified as subject to selection may be functionally important, and hence (weak) candidates for involvement in disease causation. They can also be useful in determining the adaptive differentiation of populations, and exploring hypotheses about speciation. Adaptive differentiation has traditionally been identified from differences in allele frequencies among different populations, summarised by an estimate of FST. Low outliers relative to an appropriate neutral population-genetics model indicate loci subject to balancing selection, whereas high outliers suggest adaptive (directional) selection. However, the problem of identifying statistically significant departures from neutrality is complicated by confounding effects on the distribution of FST estimates, and current methods have not yet been tested in large-scale simulation experiments. Here, we simulate data from a structured population at many unlinked, diallelic loci that are predominantly neutral but with some loci subject to adaptive or balancing selection. We develop a hierarchical-Bayesian method, implemented via Markov chain Monte Carlo (MCMC), and assess its performance in distinguishing the loci simulated under selection from the neutral loci. We also compare this performance with that of a frequentist method, based on moment-based estimates of FST. We find that both methods can identify loci subject to adaptive selection when the selection coefficient is at least five times the migration rate. Neither method could reliably distinguish loci under balancing selection in our simulations, even when the selection coefficient is twenty times the migration rate.

Adaptation, Biological↗

Multipoint linkage-disequilibrium mapping narrows location interval and identifies mutation heterogeneity.

Single-nucleotide polymorphism (SNP) genotypes were recently examined in an 890-kb region flanking the human gene CYP2D6. Single-marker and haplotype-based analyses identified, with genomewide significance (P < 10-7), a 403-kb interval displaying strong linkage disequilibrium (LD) with predicted poor-metabolizer phenotype. However, the width of this interval makes the location of causal variants difficult: for example, the interval contains seven known or predicted genes in addition to CYP2D6. We have developed the Bayesian fine-mapping software coldmap, which, applied to these genotype data, yields a 95% location interval covering only 185 kb and establishes genomewide significance for a causal locus within the region. Strikingly, our interval correctly excludes four SNPs, which individually display association with genomewide significance, including the SNP showing strongest LD (P < 10-34). In addition, coldmap distinguishes homozygous cases for the major CYP2D6 mutation from those bearing minor mutations. We further investigate a selection of SNP subsets and find that previously reported methods lead to a 38% savings in SNPs at the cost of an increase of <20% in the width of the location interval.

Bayes Theorem↗

Likelihood-based inference for genetic correlation coefficients.

We review Wright's original definitions of the genetic correlation coefficients F(ST), F(IT), and F(IS), pointing out ambiguities and the difficulties that these have generated. We also briefly survey some subsequent approaches to defining and estimating the coefficients. We then propose a general framework in which the coefficients are defined, their properties established, and likelihood-based inference implemented. Likelihood methods of inference are proposed both for bi-allelic and multi-allelic loci, within a hierarchical model which allows sharing of information both across subpopulations and across loci, but without assuming constancy in either case. This framework can be used, for example, to detect environment-related diversifying selection.

Alleles↗

Implications for DNA identification arising from an analysis of Australian forensic databases.

Previous analyses of Australian samples have suggested that populations of the same broad racial group (Caucasian, Asian, Aboriginal) tend to be genetically similar across states. This suggests that a single national Australian database for each such group may be feasible, which would greatly facilitate casework. We have investigated samples drawn from each of these groups in different Australian states, and have quantified the genetic homogeneity across states within each racial group in terms of the "coancestry coefficient" F(ST). In accord with earlier results, we find that F(ST) values, as estimated from these data, are very small for Caucasians and Asians, usually <0.5%. We find that "declared" Aborigines (which includes many with partly Aboriginal genetic heritage) are also genetically similar across states, although they display some differentiation from a "pure" Aboriginal population (almost entirely of Aboriginal genetic heritage).

Australia↗

Approximate Bayesian computation in population genetics.

We propose a new method for approximate Bayesian statistical inference on the basis of summary statistics. The method is suited to complex problems that arise in population genetics, extending ideas developed in this setting by earlier authors. Properties of the posterior distribution of a parameter, such as its mean or density curve, are approximated without explicit likelihood calculations. This is achieved by fitting a local-linear regression of simulated parameter values on simulated summary statistics, and then substituting the observed summary statistics into the regression equation. The method combines many of the advantages of Bayesian statistical inference with the computational efficiency of methods based on summary statistics. A key advantage of the method is that the nuisance parameters are automatically integrated out in the simulation step, so that the large numbers of nuisance parameters that arise in population genetics problems can be handled without difficulty. Simulation results indicate computational and statistical efficiency that compares favorably with those of alternative methods previously proposed in the literature. We also compare the relative efficiency of inferences obtained using methods based on summary statistics with those obtained directly from the data using MCMC.

Bayes Theorem↗

The DNA database search controversy.

A recent article in Biometrics (Stockmarr, 1999, 55, 671-677) has generated correspondence (56, 1274-1277; 57, 976-980) reigniting a controversy started by a 1996 report on DNA profile evidence issued by the U.S. National Research Council (NRC). The issue concerns the evidential weight of a DNA profile match when the match results from a search through a profile database. The views of both Stockmarr and the NRC report conflict with those of many statisticians working in the area, and the differing viewpoints lead to dramatically different assessments of evidence. I outline reasons why Stockmarr and the NRC report are wrong. I also briefly discuss possible reasons why forensic applications tend to be problematic for statisticians.

DNA↗