Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Haplotypic QTL mapping in an outbred pedigree.

An offspring genome can be viewed as a mosaic of chromosomal segments or haplotypes contributed by multiple founders in any quantitative trait locus (QTL) detection study but tracing these is especially complex to achieve for outbred pedigrees. QTL haplotypes can be traced from offspring back to individual founders in outbred pedigrees by combining founder-origin probabilities with fully informative flanking markers. This haplotypic method was illustrated for QTL detection using a three-generation pedigree for a woody perennial plant, Pinus taeda L. Growth rate was estimated using height measurements from ages 2 to 10 years. Using simulated and actual datasets, power of the experimental design was shown to be efficient for detecting QTLs of large effect. Using interval mapping and fully informative markers, a large QTL accounting for 11.3% of the phenotypic variance in the growth rate was detected. This same QTL was expressed at all ages for height, accounting for 7.9-12.2% of the phenotypic variance. A mixed-model inheritance was more appropriate for describing genetic architecture of growth curves in P. taeda than a strictly polygenic model. The positive QTL haplotype was traced from the offspring to its contributing founder, GP3, then the haplotypic phase for GP3 was determined by assaying haploid megagametophytes. The positive QTL haplotype was a recombinant haplotype contributed by GP3. This study illustrates the combined power of fully informative flanking markers and founder origin probabilities for (1) estimating QTL haplotype magnitude, (2) tracing founder origin and (3) determining haplotypic transmission frequency.

Haplotypes↗

Assessing optimal neural network architecture for identifying disease-associated multi-marker genotypes using a permutation test, and application to calpain 10 polymorphisms associated with diabetes.

Biallelic markers, such as single nucleotide polymorphisms (SNPs), provide greater information for localising disease loci when treated as multilocus haplotypes, but often haplotypes are not immediately available from multilocus genotypes in case-control studies. An artificial neural network allows investigation of association between disease phenotype and tightly linked markers without requiring haplotype phase and without modelling any evolutionary history for the disease-related haplotypes. The network assesses whether marker haplotypes differ between cases and controls to the extent that classification of disease status based on multi-marker genotypes is achievable. The network is "trained" to "recognise" affection status based on supplied marker genotypes, and then for each multi-marker genotype it produces outputs which aim to approximate the associated affection status. Next, the genotypes are permuted relative to affection status to produce many random datasets and the process of training and recording of outputs is repeated. The extent to which the ability to predict affection for the real dataset exceeds that for the random datasets measures the statistical significance of the association between multi-marker genotype and affection. This permutation test performs well with simulated case-control datasets, particularly when major gene effects are present. We have explored the effects of systematically varying different network parameters in order to identify their optimal values. We have applied the permutation test to 4 SNPs of the calpain 10 (CAPN10) gene typed in a case-control sample of subjects with type 2 diabetes, impaired glucose tolerance, and controls. We show that the neural network produces more highly significant evidence for association than do single marker tests corrected for the number of markers genotyped. The use of a permutation test could potentially allow conditional analyses which could incorporate known risk factors alongside marker genotypes. Permuting only the marker genotypes relative to affection status and these risk factors would allow the contribution of the markers to disease risk to be independently assessed.

Calpain↗

Evolutionary-based association analysis using haplotype data.

Association studies, both family-based and population-based, can be powerful means of detecting disease-liability alleles. To increase the information of the test, various researchers have proposed targeting haplotypes. The larger number of haplotypes, however, relative to alleles at individual loci, could decrease power because of the additional degrees of freedom required for the test. An optimal strategy would focus the test on particular haplotypes or groups of haplotypes, much as is done with cladistic-based association analysis. First suggested by Templeton et al. ([1987] Genetics 117:343-351), such analyses use the evolutionary relationships among haplotypes to produce a limited set of hypothesis tests and to increase the interpretability of these tests. To more fully utilize the information contained in the evolutionary relationships among haplotypes and in the sample, we propose generalized linear models (GLM) for the analysis of data from family-based and population-based studies. These models fully account for haplotype phase ambiguity and allow for covariates. The models are encoded into a software package (the Evolutionary-Based Haplotype Analysis Package, EHAP), which also provides for various kinds of exploratory data analysis. The exploratory analyses, such as error checking, estimation of haplotype frequencies, and tools for building cladograms, should facilitate the implementation of cladistic-based association analysis with haplotypes.

Genetic Predisposition to Disease↗

MICA-STR, HLA-B haplotypic diversity and linkage disequilibrium in the Hunan Han population of southern China.

Major histocompatibility complex (MHC) class I chain-related gene A (MICA) is located 46 kb centromeric to HLA-B and encodes a stress-inducible protein. MICA allelic variation is thought to be associated with disease susceptibility and immune response to transplants. This study was aimed to investigate the haplotypic diversity and linkage disequilibrium between human leukocyte antigen (HLA)-B and (GCT)(n) short tandem repeat in exon 5 of MICA gene (MICA-STR) in a southern Chinese Han population. Fifty-eight randomly selected nuclear families with 183 members including 85 unrelated parental samples were collected in Hunan province, southern China. HLA-B generic typing was performed by polymerase chain reaction-sequence-specific priming (PCR-SSP), and samples showing novel HLA-B-MICA-STR linkage were further typed for HLA-B allelic variation by high-resolution PCR-SSP. MICA-STR allelic variation and MICA gene deletion (MICA*Del) were detected by fluorescent PCR-size sequencing and PCR-SSP. Haplotype was determined through family segregation analysis. Statistical analysis was applied to the data of the 85 unrelated parental samples. Nineteen HLA-B specificities and seven MICA-STR allelic variants were observed in 85 unrelated parental samples, the most predominant of which were HLA-B*46, -B60, -B*13, and -B*15, and MICA*A5, MICA*A5.1 and MICA*A4, respectively. Genotype distributions of HLA-B, MICA-STR loci were consistent with Hardy-Weinberg proportions. The HLA-B-MICA-STR haplotypic phases of all 85 unrelated parental samples were unambiguously assigned, which contained 30 kinds of HLA-B, MICA-STR haplotypic combinations, nine of them have not been reported in the literature. Significant positive linkage disequilibria between certain HLA-B and MICA-STR alleles, including HLA-B*13 and MICA*A4, HLA-B*38 and MICA*A9, HLA-B*58 and MICA*A9, HLA-B*46 and MICA*A5, HLA-B*51 and MICA*A6, HLA-B*52 and MICA*A6, and HLA-B60 and MICA*A5.1, were observed. HLA-B*48 was linked to MICA*A5, MICA*A5.1 and MICA*Del. HLA-B*5801-MICA*A10 linkage was found in a family. Our data indicated a high degree of haplotypic diversity and strong linkage disequilibrium between MICA-STR and HLA-B in a southern Chinese Han population, the data will inform future studies on anthropology, donor-recipient HLA matching in clinical transplantation and HLA-linked disease association.

Asian People↗

Incorporating genotyping uncertainty in haplotype inference for single-nucleotide polymorphisms.

The accuracy of the vast amount of genotypic information generated by high-throughput genotyping technologies is crucial in haplotype analyses and linkage-disequilibrium mapping for complex diseases. To date, most automated programs lack quality measures for the allele calls; therefore, human interventions, which are both labor intensive and error prone, have to be performed. Here, we propose a novel genotype clustering algorithm, GeneScore, based on a bivariate t-mixture model, which assigns a set of probabilities for each data point belonging to the candidate genotype clusters. Furthermore, we describe an expectation-maximization (EM) algorithm for haplotype phasing, GenoSpectrum (GS)-EM, which can use probabilistic multilocus genotype matrices (called "GenoSpectrum") as inputs. Combining these two model-based algorithms, we can perform haplotype inference directly on raw readouts from a genotyping machine, such as the TaqMan assay. By using both simulated and real data sets, we demonstrate the advantages of our probabilistic approach over the current genotype scoring methods, in terms of both the accuracy of haplotype inference and the statistical power of haplotype-based association analyses.

Algorithms↗

A model-based approach to selection of tag SNPs.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are the most common type of polymorphisms found in the human genome. Effective genetic association studies require the identification of sets of tag SNPs that capture as much haplotype information as possible. Tag SNP selection is analogous to the problem of data compression in information theory. According to Shannon's framework, the optimal tag set maximizes the entropy of the tag SNPs subject to constraints on the number of SNPs. This approach requires an appropriate probabilistic model. Compared to simple measures of Linkage Disequilibrium (LD), a good model of haplotype sequences can more accurately account for LD structure. It also provides a machinery for the prediction of tagged SNPs and thereby to assess the performances of tag sets through their ability to predict larger SNP sets. RESULTS: Here, we compute the description code-lengths of SNP data for an array of models and we develop tag SNP selection methods based on these models and the strategy of entropy maximization. Using data sets from the HapMap and ENCODE projects, we show that the hidden Markov model introduced by Li and Stephens outperforms the other models in several aspects: description code-length of SNP data, information content of tag sets, and prediction of tagged SNPs. This is the first use of this model in the context of tag SNP selection. CONCLUSION: Our study provides strong evidence that the tag sets selected by our best method, based on Li and Stephens model, outperform those chosen by several existing methods. The results also suggest that information content evaluated with a good model is more sensitive for assessing the quality of a tagging set than the correct prediction rate of tagged SNPs. Besides, we show that haplotype phase uncertainty has an almost negligible impact on the ability of good tag sets to predict tagged SNPs. This justifies the selection of tag SNPs on the basis of haplotype informativeness, although genotyping studies do not directly assess haplotypes. A software that implements our approach is available.

Databases, Genetic↗

Algorithms for inferring haplotypes.

Haplotype phase information in diploid organisms provides valuable information on human evolutionary history and may lead to the development of more efficient strategies to identify genetic variants that increase susceptibility to human diseases. Molecular haplotyping methods are labor-intensive, low-throughput, and very costly. Therefore, algorithms based on formal statistical theories were shown to be very effective and cost-efficient for haplotype reconstruction. This review covers 1) population-based haplotype inference methods: Clark's algorithm, expectation-maximization (EM) algorithm, coalescence-based algorithms (pseudo-Gibbs sampler and perfect/imperfect phylogeny), and partition-ligation algorithm implemented by a fully Bayesian model (Haplotyper) or by EM (PLEM); 2) family-based haplotype inference methods; 3) the handling of genotype scoring uncertainties (i.e., genotyping errors and raw two-dimensional genotype scatterplots) in inferring haplotypes; and 4) haplotype inference methods for pooled DNA samples. The advantages and limitations of each algorithm are discussed. By using simulations based on empirical data on the G6PD gene and TNFRSF5 gene, I demonstrate that different algorithms have different degrees of sensitivity to various extents of population diversities and genotyping error rates. Future development of statistical algorithms for addressing haplotype reconstruction will resort more and more to ideas based on combinatorial mathematics, graphical models, and machine learning, and they will have profound impacts on population genetics and genetic epidemiology with the advent of the human HapMap.

Algorithms↗

The APL test: extension to general nuclear families and haplotypes and examination of its robustness.

OBJECTIVE: The Association in the Presence of Linkage test (APL) is a powerful statistical method that allows for missing parental genotypes in nuclear families. However, in its original form, the statistic does not easily extend to mixed nuclear family structures nor to multiple-marker haplotypes. Furthermore, the robustness of APL in practice has not been examined. Here we present a generalization of the APL model and examination of its robustness under a variety of non-standard scenarios. METHODS: The generalization is made possible by incorporating a bootstrap variance estimator instead of the original robust variance estimator. This allows for use of more than two affected siblings. Haplotype analysis was accomplished by combining estimation of haplotype phase into the EM algorithm. Computer simulation was used to examine robustness of the APL to departures from test assumptions. RESULTS: The extended APL tests both single-marker and multiple-marker haplotypes and shows more power than other association methods. Simulation results showed that the single-marker APL test is robust to the departure from HWE. For the haplotype test, violation of the HWE assumption can inflate type I error. We also evaluated general guidelines for the validity of APL with rare alleles and rare haplotypes. Software for the APL test is available from http://www.chg.duke.edu/research/apl.html.

Computer Simulation↗

Computation of haplotypes on SNPs subsets: advantage of the "global method".

BACKGROUND: Genetic association studies aim at finding correlations between a disease state and genetic variations such as SNPs or combinations of SNPs, termed haplotypes. Some haplotypes have a particular biological meaning such as the ones derived from SNPs located in the promoters, or the ones derived from non synonymous SNPs. All these haplotypes are "subhaplotypes" because they refer only to a part of the SNPs found in the gene. Until now, subhaplotypes were directly computed from the very SNPs chosen to constitute them, without taking into account the rest of the information corresponding to the other SNPs located in the gene. In the present work, we describe an alternative approach, called the "global method", which takes into account all the SNPs known in the region and compare the efficacy of the two "direct" and "global" methods. RESULTS: We used empirical haplotypes data sets from the GH1 promoter and the APOE gene, and 10 simulated datasets, and randomly introduced in them missing information (from 0% up to 20%) to compare the 2 methods. For each method, we used the PHASE haplotyping software since it was described to be the best. We showed that the use of the "global method" for subhaplotyping leads always to a better error rate than the classical direct haplotyping. The advantage provided by this alternative method increases with the percentage of missing genotyping data (diminution of the average error rate from 25% to less than 10%). We applied the global method software on the GRIV cohort for AIDS genetic associations and some associations previously identified through direct subhaplotyping were found to be erroneous. CONCLUSION: The global method for subhaplotyping can reduce, sometimes dramatically, the error rate on patient resolutions and haplotypes frequencies. One should thus use this method in order to minimise the risk of a false interpretation in genetic studies involving subhaplotypes. In practice the global method is always more efficient than the direct method, but a combination method taking into account the level of missing information in each subject appears to be even more interesting when the level of missing information becomes larger (>10%).

Apolipoproteins E↗

Linear parameter haplotype models with stratification.

OBJECTIVES: The question of interest is estimating the relationship between haplotypes and an outcome measure, based upon unphased genotypes. The outcome of interest might be predicting the presence of disease in a logistic model, predicting a numeric drug response in a linear model, or predicting survival time in a parametric survival model with censoring. Explanatory variables may include phased haplotype design variables, environmental variables, or interactions between them. METHODS: We extend existing generalized linear haplotype models to parametric survival outcomes. To improve the stability of model variance estimates, a profile likelihood solution is proposed. An adjustment for population stratification is also considered. Here we investigate data sampled from known 'strata' (e.g., gender or ethnicity) that influence haplotype prior probabilities and thus the regression model weights. Differing linear model variance estimates, and the effect of stratification and departures from Hardy-Weinberg Equilibrium (HWE) on parameter estimates, are compared and contrasted via simulation. RESULTS: From simulations, we observed an improvement in statistical power when using a solution to profile likelihood equations. We also saw that stratification had little impact on estimates. Haplotypes that are not in HWE had a negative impact on power to test hypotheses. Finally, profile likelihood solutions for haplotypes deviating from HWE had improved power and confidence interval coverage of regression model coefficients.

Carcinoma, Squamous Cell↗

A note on efficient computation of haplotypes via perfect phylogeny.

The problem of inferring haplotype phase from a population of genotypes has received a lot of attention recently. This is partly due to the observation that there are many regions on human genomic DNA where genetic recombination is rare (Helmuth, 2001; Daly et al., 2001; Stephens et al., 2001; Friss et al., 2001). A Haplotype Map project has been announced by NIH to identify and characterize populations in terms of these haplotypes. Recently, Gusfield introduced the perfect phylogeny haplotyping problem, as an algorithmic implication of the no-recombination in long blocks observation, together with the standard population-genetic assumption of infinite sites. Gusfield's solution based on matroid theory was followed by direct theta(nm2) solutions that use simpler techniques (Bafna et al., 2003; Eskin et al., 2003), and also bound the number of solutions to the PPH problem. In this short note, we address two questions that were left open. First, can the algorithms of Bafna et al. (2003) and Eskin et al. (2003) be sped-up to O(nm + m2) time, which would imply an O(nm) time-bound for the PPH problem? Second, if there are multiple solutions, can we find one that is most parsimonious in terms of the number of distinct haplotypes. We give reductions that suggests that the answer to both questions is "no." For the first problem, we show that computing the output of the first step (in either method) is equivalent to Boolean matrix multiplication. Therefore, the best bound we can presently achieve is O(nm(omega-1)), where omega < or = 2.52 is the exponent of matrix multiplication. Thus, any linear time solution to the PPH problem likely requires a different approach. For the second problem of computing a PPH solution that minimizes the number of distinct haplotypes, we show that the problem is NP-hard using a reduction from Vertex Cover (Garey and Johnson, 1979).

Computational Biology↗

Multilocus linkage disequilibrium mapping by the decay of haplotype sharing with samples of related individuals.

We consider the problem of multilocus linkage disequilibrium (LD) mapping of a trait-associated variant from case-control samples in which some individuals may be related. Our method, which we call DHS-R, is an extension of the decay of haplotype sharing (DHS) method of McPeek and Strahs and Strahs and McPeek. The DHS-R method shares the main features of the DHS method: (1) it allows construction of a confidence interval for the location of a trait-associated variant; (2) it allows for missing observations and unphased genotype data, with the uncertainty in the haplotypes taken into account in the analysis; and (3) it allows for heterogeneity, mutation, recombination, and background LD. The main advances of the DHS-R are (1) the ability to include individuals of arbitrary known relationship (including inbreeding) in the case and control samples; (2) an extension to allow partially-phased haplotypes derived from case-parent trio genotype data; and (3) an extension to allow for genotyping error in the model. Our method, which uses a hidden Markov model for likelihood calculation and maximization, has the advantage of being computationally feasible even in a large, complex pedigree. Simulations based on a 13-generation, 1,623-member Hutterite pedigree demonstrate accurate coverage of the confidence intervals for location of the variant. We apply the method to fine-mapping of a susceptibility locus for bronchial hyperresponsiveness (BHR) in the Hutterites. The results confirm the importance of taking into account the relatedness of individuals in LD mapping.

Algorithms↗

A candidate gene association study on preterm delivery: application of high-throughput genotyping technology and advanced statistical methods.

Preterm delivery (PTD) is the leading cause of infant mortality and morbidity worldwide. The etiology of PTD is largely unknown but is believed to be complex, encompassing multiple genetic and environmental determinants. To date, reports of genetic studies on PTD are sparse. We conducted a large-scale case-control study exploring the associations of 426 single-nucleotide polymorphisms with PTD in 300 mothers with PTD and 458 mothers with term deliveries at the Boston Medical Center. Twenty-five candidate genes were included in the final haplotype analysis, and a significant association of F5 gene haplotype with PTD was revealed and remained significant after Bonferroni correction for multiple testing (P=0.025). We applied different statistical algorithms (both Gibbs sampling and expectation-maximization) in reconstructing haplotype phases and different tests (both likelihood ratio test and permutation test) in association analyses, and all yielded similar results. We also performed exploratory ethnicity-specific analyses, which confirmed the consistent findings of the F5 gene across the ethnic groups. Moreover, IL1R2 (P=0.002 in Blacks), NOS2A (P<0.001 in Whites) and OPRM1 (P=0.004 in Hispanics) gene haplotypes were associated with PTD in specific ethnic groups but not at global significance level. In summary, our results underscore the potentially important role of F5 gene variants in the pathogenesis of PTD, and demonstrate the utility of high-throughput genotyping and a haplotype-based approach in dissecting genetic basis of complex traits.

Algorithms↗

The evolution of separate sexes in waterhemp is associated with surprising chromosomal diversity and complexity.

The evolution of separate sexes is hypothesized to occur through distinct pathways involving few large-effect or many small-effect alleles. However, we lack empirical evidence for how these different genetic architectures shape the transition from quantitative variation in sex expression to distinct male and female phenotypes. To explore these processes, we leveraged the recent transition of Amaranthus tuberculatus to dioecy within a predominantly monoecious genus, along with a sex-phenotyped population genomic dataset, and six newly generated chromosome-level haplotype phased assemblies. We identify a ~3&#x2009;Mb region strongly associated with sex through complementary SNP genotype and sequence-depth-based analyses. Comparative genomics of these proto-sex chromosomes within the species and across the Amaranthus genus demonstrates remarkable variability in their structure and genic content, including numerous polymorphic inversions. No such inversion underlies the extended linkage we observe associated with sex determination. Instead, we identify a complex presence/absence polymorphism reflecting substantial Y-haplotype variation-structured by ancestry, geography, and habitat-but only partially explaining phenotyped sex. Just over 10% of sexed individuals show phenotype-genotype mismatch in the sex-linked region, and along with observation of leakiness in the phenotypic expression of sex, suggest additional modifiers of sex and dynamic gene content within and between the proto-X and Y. Together, this work reveals a complex genetic architecture of sex determination in A. tuberculatus characterized by the maintenance of substantial haplotype diversity, and variation in the expression of sex.

Haplotypes↗

Haplotype information and linkage disequilibrium mapping for single nucleotide polymorphisms.

Single nucleotide polymorphisms in the human genome have become an increasingly popular topic in that their analyses promise to be a key step toward personalized medicine. We investigate two related questions, how much the haplotype information contributes to linkage disequilibrium (LD) mapping and whether an in silico haplotype construction preceding the LD analysis can help. For disease gene mapping, using both simulated and real data sets on cystic fibrosis and the Alzheimer disease, we reached the following conclusions: (1) for simple Mendelian diseases, in which case a tractable full statistical model can be developed, the loss of haplotype information for either control or disease data do not have a great impact on LD fine mapping, and haplotype inference should be carried out jointly with LD mapping; (2) for complex diseases, inferring haplotype phases for individuals prior to LD mapping helps achieve a better accuracy. An improved version of the linkage disequilibrium mapping program, BLADE v2, is available at http://www.fas.harvard.edu/junliu/TechRept/03folder/bladev2.tgz.

Algorithms↗

Robust testing of haplotype/disease association.

Haplotypes, the combination of closely linked alleles that fall on the same chromosome, show great promise for studying the genetic components of complex diseases. However, when only multilocus genotype data are available, statistical approaches need to be employed to resolve haplotype phase ambiguity. Recently, we have proposed an approach to testing and estimating haplotype/disease association that is invariant to any existing genetic structure in the population. Here we evaluate this approach by applying it to the Genetic Analysis Workshop 14 simulated data.

Bias↗

Contrasting multi-site genotypic distributions among discordant quantitative phenotypes: the APOA1/C3/A4/A5 gene cluster and cardiovascular disease risk factors.

Most tests of association between DNA sequence variation and quantitative phenotypes in samples of randomly chosen individuals rely on specification of genotypic strata followed by comparison of phenotypes across these strata. This strategy often succeeds when phenotypic differences are caused by one or two single nucleotide polymorphisms (SNPs) among the surveyed markers. However, when multiple-SNP haplotypes account for observed phenotypic variation, identification of the best partitioning requires examination of an inordinate number of SNP combinations. An alternative approach is to rank individuals by their phenotypic measures and ask whether attributes of the genotypic variation show a non-random distribution along this phenotypic ranking. One simple version of this strategy selects the top and bottom tails of the distribution, and then tests whether genotypes from these two samples are drawn from a single population. This framework does not require the recovery of phased haplotypes and allows contrasts between large numbers of sites at once. We use a method based on this approach to identify associations between plasma triglyceride level, a risk factor for cardiovascular disease, and multi-site genotypes located in the APOA1/C3/A4/A5 cluster of apolipoprotein genes in unrelated individuals (1,071 African-American females, 780 African-American males, 1,036 European-American females, and 930 European-American males) sampled from four US cities as part of the Coronary Artery Risk Development in Young Adults (CARDIA) study. Method performance is investigated using simulations that model genealogical variation and different genetic architectures. Results indicate that this multi-site test can identify genotype-phenotype associations with reasonable power, including those generated by some simple epistatic models.

Adult↗

Contrasting linkage-disequilibrium patterns between cases and controls as a novel association-mapping method.

Identification and description of genetic variation underlying disease susceptibility, efficacy, and adverse reactions to drugs remains a difficult problem. One of the important steps in the analysis of variation in a candidate region is the characterization of linkage disequilibrium (LD). In a region of genetic association, the extent of LD varies between the case and the control groups. Separate plots of pairwise standardized measures of LD (e.g., D') for cases and controls are often presented for a candidate region, to graphically convey case-control differences in LD. However, the observed graphic differences lack statistical support. Therefore, we suggest the "LD contrast" test to compare whole matrices of disequilibrium between two samples. A common technique of assessing LD when the haplotype phase is unobserved is the expectation-maximization algorithm, with the likelihood incorporating the assumption of Hardy-Weinberg equilibrium (HWE). This approach presents a potential problem in that, in the region of genetic association, the HWE assumption may not hold when samples are selected on the basis of phenotypes. Here, we present a computationally feasible approach that does not assume HWE, along with graphic displays and a statistical comparison of pairwise matrices of LD between case and control samples. LD-contrast tests provide a useful addition to existing tools of finding and characterizing genetic associations. Although haplotype association tests are expected to provide superior power when susceptibilities are primarily determined by haplotypes, the LD-contrast tests demonstrate substantially higher power under certain haplotype-driven disease models.

Case-Control Studies↗