Search PubMed⌕ Search

Biomedical subjects

N L Kaplan

Publications and source records attributed to N L Kaplan.

At least 19 recordsLinked to original sources

Efficient use of siblings in testing for linkage and association.

Tests of linkage and association between a disease and either a candidate gene or marker allele can be based on sibships with at least one affected and one unaffected sibling. However, specialized techniques are required to account for within-sibship correlation if some sibships contain more than one affected or more than one unaffected sib. In this paper, we propose Within Sibship Paired Resampling (WSPR), a technique that is designed to test the null hypothesis of no linkage or no association, even when sibships contain variable numbers of sibs. One repeatedly generates data subsets based on randomly sampling one affected and one unaffected sibling from each sibship, and each subset is analyzed individually. Then, evidence is combined by averaging results across these resampled data sets, applying a variance expression that implicitly accounts for the correlation among siblings. While the general WSPR procedure allows for numerous testing strategies, we describe two in detail. Simulation results for scenarios with varying degrees of population stratification demonstrate good power for the WSPR testing methods compared to the sib TDT (S-TDT) and the sibship disequilibrium test (SDT).

Computer Simulation↗

Analysis of single nucleotide polymorphisms in candidate genes using the pedigree disequilibrium test.

The pedigree disequilibrium test (PDT) has been proposed recently as a test for association in general pedigrees [Martin et al., Am J Hum Genet 67:146-54, 2000]. The Genetic Analysis Workshop (GAW) 12 simulated data, with many extended pedigrees, is an example the type of data to which the PDT is ideally suited. In replicate 42 from the general population the PDT correctly identifies candidate genes 1, 2, and 6 as containing single nucleotide polymorphisms (SNPs) that are significantly associated with the disease. We also applied the truncated product method (TPM) [Zaykin et al., Genet Epidemiol, in press] to combine p-values in overlapping windows across the genes. Our results show that the TPM is helpful in identifying significant SNPs as well as removing spurious false positives. Our results indicate that, using the PDT, functional disease-associated SNPs can be successfully identified with a dense map of moderately polymorphic SNPs.

Genetic Predisposition to Disease↗

Power calculations for a general class of tests of linkage and association that use nuclear families with affected and unaffected sibs.

Family-based tests of association are now often used when trying to fine-map a disease susceptibility locus. Recently, several tests of linkage and association have been proposed that use nuclear families with multiple affected and unaffected sibs rather than just case-parent triads. In this paper we propose a test that generalizes these previous tests. Formulae are derived to calculate the power of the test for a randomly mating population. These power calculations are used to determine conditions under which it is advantageous to include unaffected sibs in the analysis.

Alleles↗

A test for linkage and association in general pedigrees: the pedigree disequilibrium test.

Family-based tests of linkage disequilibrium typically are based on nuclear-family data including affected individuals and their parents or their unaffected siblings. A limitation of such tests is that they generally are not valid tests of association when data from related nuclear families from larger pedigrees are used. Standard methods require selection of a single nuclear family from any extended pedigrees when testing for linkage disequilibrium. Often data are available for larger pedigrees, and it would be desirable to have a valid test of linkage disequilibrium that can use all potentially informative data. In this study, we present the pedigree disequilibrium test (PDT) for analysis of linkage disequilibrium in general pedigrees. The PDT can use data from related nuclear families from extended pedigrees and is valid even when there is population substructure. Using computer simulations, we demonstrated validity of the test when the asymptotic distribution is used to assess the significance, and examined statistical power. Power simulations demonstrate that, when extended pedigree data are available, substantial gains in power can be attained by use of the PDT rather than existing methods that use only a subset of the data. Furthermore, the PDT remains more powerful even when there is misclassification of unaffected individuals. Our simulations suggest that there may be advantages to using the PDT even if the data consist of independent families without extended family information. Thus, the PDT provides a general test of linkage disequilibrium that can be widely applied to different data structures.

Alleles↗

A Monte Carlo procedure for two-stage tests with correlated data.

One strategy for mapping disease loci using marker-disease associations is to test for association with case-control samples and follow up a positive result with a family-based test. Using a family-based test in the second stage can help provide protection against false-positive results that can result from use of inappropriate controls and provides assurance that association identified in the first stage is occurring between linked loci. It is crucial for this two-stage strategy that the first stage be as powerful as possible to detect association since only positive results are tested in the second stage. In certain situations, the power of the first-stage test can be increased by combining the case-control and family data. However, this introduces correlation between the first- and second-stage tests, and treating them as independent tests causes a bias. Here we propose a Monte Carlo method that accounts for the correlation and provides the correct significance level for the second-stage test. We also discuss the use of a two-stage procedure when doing a genome scan for the data presented in the Genetic Analysis Workshop 9 study.

Genetic Markers↗

Circumventing multiple testing: a multilocus Monte Carlo approach to testing for association.

Advances in marker technology have made a dense marker map a reality. If each marker is considered separately, and separate tests for association with a disease gene are performed, then multiple testing becomes an issue. A common solution uses a Bonferroni correction to account for multiple tests performed. However, with dense marker maps, neighboring markers are tightly linked and may have associated alleles; thus tests at nearby marker loci may not be independent. When alleles at different marker loci are associated, the Bonferroni correction may lead to a conservative test, and hence a power loss. As an alternative, for tests of association that use family data, we propose a Monte Carlo procedure that provides a global assessment of significance. We examine the case of tightly linked markers with varying amounts of association between them. Using computer simulations, we study a family-based test for association (the transmission/disequilibrium test), and compare its power when either the Bonferroni or Monte Carlo procedure is used to determine significance. Our results show that when the alleles at different marker loci are not associated, using either procedure results in tests with similar power. However, when alleles at linked markers are associated, the test using the Monte Carlo procedure is more powerful than the test using the Bonferroni procedure. This proposed Monte Carlo procedure can be applied whenever it is suspected that markers examined have high amounts of association, or as a general approach to ensure appropriate significance levels and optimal power.

Alleles↗

Removing the sampling restrictions from family-based tests of association for a quantitative-trait locus.

One strategy for localization of a quantitative-trait locus (QTL) is to test whether the distribution of a quantitative trait depends on the number of copies of a specific genetic-marker allele that an individual possesses. This approach tests for association between alleles at the marker and the QTL, and it assumes that association is a consequence of the marker being physically close to the QTL. However, problems can occur when data are not from a homogeneous population, since associations can arise irrespective of a genetic marker being in physical proximity to the QTL-that is, no information is gained regarding localization. Methods to address this problem have recently been proposed. These proposed methods use family data for indirect stratification of a population, thereby removing the effect of associations that are due to unknown population substructure. They are, however, restricted in terms of the number of children per family that can be used in the analysis. Here we introduce tests that can be used on family data with parent and child genotypes, with child genotypes only, or with a combination of these types of families, without size restrictions. Furthermore, equations that allow one to determine the sample size needed to achieve desired power are derived. By means of simulation, we demonstrate that the existing tests have an elevated false-positive rate when the size restrictions are not followed and that a good deal of information is lost as a result of adherence to the size restrictions. Finally, we introduce permutation procedures that are recommended for small samples but that can also be used for extensions of the tests to multiallelic markers and to the simultaneous use of more than one marker.

Alleles↗

Two tests of association for a susceptibility locus for families of variable size: an example using two sampling strategies.

A two-stage approach was used to analyze Problem 2 simulated data from Genetic Analysis Workshop 11. In the first stage, we tested for linkage with the Haseman-Elston test in SIBPAL. Markers that were significant in the first stage were followed up with two types of association tests. These association tests differ in the type of family information used: 1) parental transmissions to affected children or 2) differences in marker allele frequencies between affected and unaffected siblings. We also explored how the conclusions changed when different sampling strategies were used. Of particular interest was whether the entire data set should be used to test for both linkage and association or whether the data set should be halved to allow for replication of the initial association results.

Environment↗

Marker selection for the transmission/disequilibrium test, in recently admixed populations.

Recent admixture between genetically differentiated populations can result in high levels of association between alleles at loci that are <=10 cM apart. The transmission/disequilibrium test (TDT) proposed by Spielman et al. (1993) can be a powerful test of linkage between disease and marker loci in the presence of association and therefore could be a useful test of linkage in admixed populations. The degree of association between alleles at two loci depends on the differences in allele frequencies, at the two loci, in the founding populations; therefore, the choice of marker is important. For a multiallelic marker, one strategy that may improve the power of the TDT is to group marker alleles within a locus, on the basis of information about the founding populations and the admixed population, thereby collapsing the marker into one with fewer alleles. We have examined the consequences of collapsing a microsatellite into a two-allele marker, when two founding populations are assumed for the admixed population, and have found that if there is random mating in the admixed population, then typically there is a collapsing for which the power of the TDT is greater than that for the original microsatellite marker. A method is presented for finding the optimal collapsing that has minimal dependence on the disease and that uses estimates either of marker allele frequencies in the two founding populations or of marker allele frequencies in the current, admixed population and in one of the founding populations. Furthermore, this optimal collapsing is not always the collapsing with the largest difference in allele frequencies in the founding populations. To demonstrate this strategy, we considered a recent data set, published previously, that provides frequency estimates for 30 microsatellites in 13 populations.

Alleles↗

A comparative study of sibship tests of linkage and/or association.

Population-based tests of association have used data from either case-control studies or studies based on trios (affected child and parents). Case-control studies are more prone to false-positive results caused by inappropriate controls, which can occur if, for example, there is population admixture or stratification. An advantage of family-based tests is that cases and controls are well matched, but parental data may not always be available, especially for late-onset diseases. Three recent family-based tests of association and linkage utilize unaffected siblings as surrogates for untyped parents. In this paper, we propose an extension of one of these tests. We describe and compare the four tests in the context of a complex disease for both biallelic and multiallelic markers, as well as for sibships of different sizes. We also examine the consequences of having some parental data in the sample.

Alleles↗

A Monte Carlo permutation approach to choosing an affection status model for bipolar affective disorder.

A permutation test is proposed for assessing affection status models. The test uses marker data from regions with prior evidence of linkage to susceptibility genes, and three different test statistics are examined. We applied the test to the GAW10 data and found no evidence on chromosome 18 to reject the affection status model that groups individuals diagnosed with either bipolar I, bipolar II or unipolar. The chromosome 5 data gave similar results, and further suggested that individuals diagnosed with unipolar-single episode not be included as affected. A preliminary power study suggested that one of the proposed statistics, S, is to be preferred in certain circumstances.

Bipolar Disorder↗

Tests for linkage and association in nuclear families.

The transmission/disequilibrium test (TDT) originally was introduced to test for linkage between a genetic marker and a disease-susceptibility locus, in the presence of association. Recently, the TDT has been used to test for association in the presence of linkage. The motivation for this is that linkage analysis typically identifies large candidate regions, and further refinement is necessary before a search for the disease gene is begun, on the molecular level. Evidence of association and linkage may indicate which markers in the region are closest to a disease locus. As a test of linkage, transmissions from heterozygous parents to all of their affected children can be included in the TDT; however, the TDT is a valid chi2 test of association only if transmissions to unrelated affected children are used in the analysis. If the sample contains independent nuclear families with multiple affected children, then one procedure that has been used to test for association is to select randomly a single affected child from each sibship and to apply the TDT to those data. As an alternative, we propose two statistics that use data from all of the affected children. The statistics give valid chi2 tests of the null hypothesis of no association or no linkage and generally are more powerful than the TDT with a single, randomly chosen, affected child from each family.

Alleles↗

Power studies for the transmission/disequilibrium tests with multiple alleles.

Case-control studies compare marker-allele distributions in affected and unaffected individuals, and significant results suggest linkage but may simply reflect population structure. For markers with m alleles (m > or = 2), a McNemar-like statistic, I, estimates the level of population association between marker and disease loci. To test for linkage after significant case-control tests, within-family tests are performed. These operate on the contingency table, with i, jth element equal to the number of parents that transmit marker allele Mi and do not transmit marker allele Mi to an affected offspring. The dimension of the table is the number of alleles at the marker locus. Three test statistics have recently been proposed in the literature: Tc compares symmetric pairs of cells (i, j) and (j, i), Tm compares row and column totals for the same marker allele, and a likelihood ratio statistic Tl uses all the cells in the table. In addition, we consider a new statistic, Tmhet, that uses only the heterozygous parents and is approximately chi2 with (m - 1) df. We use a Monte Carlo test to guarantee valid tests and to demonstrate the inferiority of Tc and the equality of Tm and Tl in terms of power. The power of the Tmhet test is close but not always equal to the power of the Tm test. We also show that under the alternative hypothesis of linkage, Tm is approximately noncentral chi2 with (m - 1) df and noncentrality parameter 2NT(1 - 2theta)2I*, when data on single affecteds in NT families are used. If the disease has a low population frequency, then I* is estimated using the case-control statistic I. This offers a basis for choosing sample size, or choosing a marker system.

Alleles↗

The coalescent process and background selection.

Some statistical properties of gene trees are described for a model with background deleterious mutations. It is argued that the history of a small sample of genes under this model is a continuous time Markov chain that quickly reaches stationarity. This observation leads to simple expressions for the expected nucleotide diversity and suggests that the frequency spectrum in small samples should be approximately the same as under a strict neutral model. The results concerning expected nucleotide diversity are compared with observed variation on the third chromosome of Drosophila melanogaster.

Animals↗

The hitchhiking effect on the site frequency spectrum of DNA polymorphisms.

The level of DNA sequence variation is reduced in regions of the Drosophila melanogaster genome where the rate of crossing over per physical distance is also reduced. This observation has been interpreted as support for the simple model of genetic hitchhiking, in which directional selection on rare variants, e.g., newly arising advantageous mutants, sweeps linked neutral alleles to fixation, thus eliminating polymorphisms near the selected site. However, the frequency spectra of segregating sites of several loci from some populations exhibiting reduced levels of nucleotide diversity and reduced numbers of segregating sites did not appear different from what would be expected under a neutral equilibrium model. Specifically, a skew toward an excess of rare sites was not observed in these samples, as measured by Tajima's D. Because this skew was predicted by a simple hitchhiking model, yet it had never been expressed quantitatively and compared directly to DNA polymorphism data, this paper investigates the hitchhiking effect on the site frequency spectrum, as measured by Tajima's D and several other statistics, using a computer simulation model based on the coalescent process and recurrent hitchhiking events. The results presented here demonstrate that under the simple hitchhiking model (1) the expected value of Tajima's D is large and negative (indicating a skew toward rare variants), (2) that Tajima's test has reasonable power to detect a skew in the frequency spectrum for parameters comparable to those from actual data sets, and (3) that the Tajima's Ds observed in several data sets are very unlikely to have been the result of simple hitchhiking. Consequently, the simple hitchhiking model is not a sufficient explanation for the DNA polymorphism at those loci exhibiting a decreased number of segregating sites yet not exhibiting a skew in the frequency spectrum.

Animals↗

Deleterious background selection with recombination.

An analytic expression for the expected nucleotide diversity is obtained for a neutral locus in a region with deleterious mutation and recombination. Our analytic results are used to predict levels of variation for the entire third chromosome of Drosophila melanogaster. The predictions are consistent with the low levels of variation that have been observed at loci near the centromeres of the third chromosome of D. melanogaster. However, the low levels of variation observed near the tips of this chromosome are not predicted using currently available estimates of the deleterious mutation rate and of selection coefficients. If considerably smaller selection coefficients are assumed, the low observed levels of variation at the tips of the third chromosome are consistent with the background selection model.

Animals↗

Likelihood methods for locating disease genes in nonequilibrium populations.

Until recently, attempts to map disease genes on the basis of population associations with linked markers have been based on expected values of linkage disequilibrium. These methods suffer from the large variances imposed on disequilibrium measures by the evolutionary process, but a more serious problem for many diseases is that they assume an equilibrium population. For diseases that arose only a few hundred generations ago, it is more appropriate to concentrate on the initial growth phase of the disease. We invoke a Poisson branching process for this early growth, and estimate the likelihood for the recombination fraction between marker and disease loci, on the basis of simulated disease populations. The limits of the resulting support intervals for the recombination fraction vary inversely with the age of the disease in generations. We illustrate the procedure with data on cystic fibrosis and diastrophic dysplasia, for which the method appears appropriate, and for Friedreich ataxia and Huntington disease, for which it does not. A valuable aspect of the method is the ability in some cases to compare likelihoods of the three orders for a disease locus and two linked marker loci.

Chromosome Mapping↗