Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Haplotype-based quantitative trait mapping using a clustering algorithm.

BACKGROUND: With the availability of large-scale, high-density single-nucleotide polymorphism (SNP) markers, substantial effort has been made in identifying disease-causing genes using linkage disequilibrium (LD) mapping by haplotype analysis of unrelated individuals. In addition to complex diseases, many continuously distributed quantitative traits are of primary clinical and health significance. However the development of association mapping methods using unrelated individuals for quantitative traits has received relatively less attention. RESULTS: We recently developed an association mapping method for complex diseases by mining the sharing of haplotype segments (i.e., phased genotype pairs) in affected individuals that are rarely present in normal individuals. In this paper, we extend our previous work to address the problem of quantitative trait mapping from unrelated individuals. The method is non-parametric in nature, and statistical significance can be obtained by a permutation test. It can also be incorporated into the one-way ANCOVA (analysis of covariance) framework so that other factors and covariates can be easily incorporated. The effectiveness of the approach is demonstrated by extensive experimental studies using both simulated and real data sets. The results show that our haplotype-based approach is more robust than two statistical methods based on single markers: a single SNP association test (SSA) and the Mann-Whitney U-test (MWU). The algorithm has been incorporated into our existing software package called HapMiner, which is available from our website at http://www.eecs.case.edu/~jxl175/HapMiner.html. CONCLUSION: For QTL (quantitative trait loci) fine mapping, to identify QTNs (quantitative trait nucleotides) with realistic effects (the contribution of each QTN less than 10% of total variance of the trait), large samples sizes (>or= 500) are needed for all the methods. The overall performance of HapMiner is better than that of the other two methods. Its effectiveness further depends on other factors such as recombination rates and the density of typed SNPs. Haplotype-based methods might provide higher power than methods based on a single SNP when using tag SNPs selected from a small number of samples or some other sources (such as HapMap data). Rank-based statistics usually have much lower power, as shown in our study.

Algorithms↗

Absence of alpha-1 antitrypsin deficiency alleles (S and Z) in Japanese and Korean patients with aneurysmal subarachnoid hemorrhage.

BACKGROUND AND PURPOSE: A possible association has been proposed for the formation of intracranial aneurysm (IA) and deficiency alleles (S and Z) of the alpha1-antitrypsin (AAT) gene. We extensively screened this gene in Japanese and Korean patients with aneurysmal subarachnoid hemorrhage. METHODS: Seven allelic variants, including S and Z alleles, were genotyped by direct sequencing of genomic DNA obtained from 195 and 189 ruptured IA patients and 195 and 94 controls in Japanese and Koreans, respectively. The haplotype in phase-unknown samples was constructed with the expectation-maximization method. Differences in allelic frequencies between patients and controls were evaluated by Fisher exact test. RESULTS: No significant differences in allelic frequencies were observed at all 7 variants between ruptured IA patients and controls. We could not detect the S and Z alleles of the AAT gene in Japanese and Korean populations. CONCLUSIONS: AAT deficiency may not be a common genetic risk factor for aneurysmal subarachnoid hemorrhage in Japanese and Koreans.

Asian People↗

A totally synthetic polyoxime malaria vaccine containing Plasmodium falciparum B cell and universal T cell epitopes elicits immune responses in volunteers of diverse HLA types.

This open-labeled phase I study provides the first demonstration of the immunogenicity of a precisely defined synthetic polyoxime malaria vaccine in volunteers of diverse HLA types. The polyoxime, designated (T1BT(*))(4)-P3C, was constructed by chemoselective ligation, via oxime bonds, of a tetrabranched core with a peptide module containing B cell epitopes and a universal T cell epitope of the Plasmodium falciparum circumsporozoite protein. The triepitope polyoxime malaria vaccine was immunogenic in the absence of any exogenous adjuvant, using instead a core modified with the lipopeptide P3C as an endogenous adjuvant. This totally synthetic vaccine formulation can be characterized by mass spectroscopy, thus enabling the reproducible production of precisely defined vaccines for human use. The majority of the polyoxime-immunized volunteers (7/10) developed high levels of anti-repeat Abs that reacted with the native circumsporozoite on P. falciparum sporozoites. In addition, these seven volunteers all developed T cells specific for the universal epitope, termed T(*), which was originally defined using CD4(+) T cells from protected volunteers immunized with irradiated P. falciparum sporozoites. The excellent correlation of T(*)-specific cellular responses with high anti-repeat Ab titers suggests that the T(*) epitope functioned as a universal Th cell epitope, as predicted by previous peptide/HLA binding assays and by immunogenicity studies in mice of diverse H-2 haplotypes. The current phase I trial suggests that polyoximes may prove useful for the development of highly immunogenic, multicomponent synthetic vaccines for malaria, as well as for other pathogens.

Adult↗

Phosphodiesterase 4D gene, ischemic stroke, and asymptomatic carotid atherosclerosis.

BACKGROUND AND PURPOSE: Phosphodiesterase 4D (PDE4D) was identified recently as the first novel stroke gene to predispose to ischemic stroke independently of conventional risk factors. An association was only found with large vessel and cardioembolic stroke, suggesting a mechanism of accelerated atherosclerosis. We sought to replicate this association in ischemic stroke as a whole, and individual stroke subtypes, in a non-Icelandic European population. To assess a role in early atherosclerosis, we also sought associations with underlying asymptomatic atherosclerosis itself, assessed by carotid ultrasound in a community population. METHODS: A total of 737 consecutive white patients with stroke and 933 white community controls free of symptomatic cerebrovascular disease were examined using a case control methodology. For association with atherosclerosis, intima-media thickness (IMT) in a community population (n=1000) was assessed using carotid ultrasound. Nineteen single nucleotide polymorphisms (SNPs) and 1 minisatellite in the PDE4D gene were determined, with haplotyping undertaken using Phase 2.0. RESULTS: No association with ischemic stroke overall was identified. Six of the 19 SNPs were associated with cardioembolic stroke and 2 different SNPs with large vessel disease. There was no association with carotid artery IMT or carotid plaque in the asymptomatic community subjects. CONCLUSIONS: The PDE4D gene is not a major risk factor for ischemic stroke, or early atherosclerosis, within the 2 European population samples studied. On analysis of individual stroke subtypes, there is a possible association with cardioembolic stroke, but the lack of association with carotid IMT and plaque would suggest that this is via a mechanism other than accelerated atherosclerosis.

3',5'-Cyclic-AMP Phosphodiesterases↗

Quantifying the amount of missing information in genetic association studies.

Many genetic analyses are done with incomplete information; for example, unknown phase in haplotype-based association studies. Measures of the amount of available information can be used for efficient planning of studies and/or analyses. In particular, the linkage disequilibrium (LD) between two sets of markers can be interpreted as the amount of information one set of markers contains for testing allele frequency differences in the second set, and measuring LD can be viewed as quantifying information in a missing data problem. We introduce a framework for measuring the association between two sets of variables; for example, genotype data for two distinct groups of markers, or haplotype and genotype data for a given set of polymorphisms. The goal is to quantify how much information is in one data set, e.g. genotype data for a set of SNPs, for estimating parameters that are functions of frequencies in the second data set, e.g. haplotype frequencies, relative to the ideal case of actually observing the complete data, e.g. haplotypes. In the case of genotype data on two mutually exclusive sets of markers, the measure determines the amount of multi-locus LD, and is equal to the classical measure r(2), if the sets consist each of one bi-allelic marker. In general, the measures are interpreted as the asymptotic ratio of sample sizes necessary to achieve the same power in case-control testing. The focus of this paper is on case-control allele/haplotype tests, but the framework can be extended easily to other settings like regressing quantitative traits on allele/haplotype counts, or tests on genotypes or diplotypes. We highlight applications of the approach, including tools for navigating the HapMap database [The International HapMap Consortium, 2003], and genotyping strategies for positional cloning studies.

Alleles↗

Interleukin-10 promoter microsatellite polymorphisms in systemic lupus erythematosus: association with the anti-Sm immune response.

OBJECTIVES: Overproduction of interleukin-10 (IL-10) is a pivotal feature in the pathophysiology of systemic lupus erythematosus (SLE). In vitro IL-10 secretion has previously been related to haplotypes of the IL-10 promoter microsatellite polymorphisms IL10.R and IL10.G. Published data concerning the association of IL10.G alleles with susceptibility to SLE are inconsistent in different ethnic populations. We analysed the association of IL-10 promoter microsatellite polymorphisms with disease susceptibility and manifestations in German Caucasian patients with SLE. METHODS: Two hundred and ten (210) SLE patients fulfilling the 1997 revised ACR criteria and 158 ethnically, age- and sex-matched healthy controls were genotyped for the IL-10 promoter microsatellite polymorphisms by fragment length analysis. Haplotypes were reconstructed using a Bayesian coalescent theory-based method with PHASE software. Allele and haplotype distributions were compared between patients and controls and between subgroups of patients with different clinical and immunopathological findings. RESULTS: In the study population no significant associations of individual IL10.R and G alleles or their haplotypes with susceptibility to SLE or major clinical manifestations were observed. By contrast, alleles G14 and G15 and haplotypes R2-G14 and R2-G15 were significantly over-represented in anti-Sm antibody-positive patients. CONCLUSIONS: The IL-10 promoter microsatellite polymorphisms and their haplotypes do not constitute a major risk factor for SLE in German Caucasians. However, the identification of genetic markers such as the IL-10 high-response haplotype R2-G14 predisposing for the production of anti-Sm antibodies may help to elucidate the conditions that lead to the development of SLE.

Adolescent↗

Haplotypes in the tumour necrosis factor region and myeloma.

This study described the haplotypic structure across a region of chromosome 6 including the tumour necrosis factor (TNF) gene, and investigated its influence on the aetiology of myeloma. A total of 181 myeloma cases from the Medical Research Council Myeloma VII trial and 233 controls from the Leukaemia Research Fund Case Control Study of Adult Acute Leukaemia were included in the analysis. Genotyping by induced heteroduplex generator analysis was carried out for single nucleotide polymorphisms (SNP) located at positions -1031, -863, -857, -308 and -238 of the 5' promoter region of TNF-alpha gene, and 252 in the LT-alpha gene; and five microsatellites, TNFa, b, c, d and e. Haplotypes were inferred statistically using the phase algorithm. A limited diversity of haplotypes was observed, with the majority of variation described by 12 frequent haplotypes. Detailed characterization of the haplotype did not provide greater determination of disease risk beyond that described by the TNF-alpha-308 SNP. Some evidence was provided for a decreased risk of myeloma associated with the TNF-alpha-308 variant allele A, odds ratio, 0.57; 95% confidence interval, 0.38-0.86. The results of this study did not support our starting hypothesis; that high producer haplotypes at the TNF locus are associated with an increased risk of developing myeloma.

Adult↗

Integration of HapMap-based SNP pattern analysis and gene expression profiling reveals common SNP profiles for cancer therapy outcome predictor genes.

Recent completion of the initial phase of a haplotype map of human genome (www.hapmap.org) provides opportunity for integrative analysis on a genome-wide scale of microarray-based gene expression profiling and SNP variation patterns for discovery of cancer-causing genes and genetic markers of therapy outcome. Here we applied this approach for analysis of SNPs of cancer-associated genes, expression profiles of which predicts the likelihood of treatment failure and death after therapy in patients diagnosed with multiple types of cancer. Unexpectedly, this analysis reveals a common SNP pattern for a majority (60 of 74; 81%) of analyzed cancer treatment outcome predictor (CTOP) genes. Our analysis suggests that heritable germ-line genetic variations driven by geographically localized form of natural selection determining population differentiations may have a significant impact on cancer treatment outcome by influencing the individual's gene expression profile. We demonstrate a translational utility of this approach by building a highly informative CTOP algorithm combining prognostic power of multiple gene expression-based CTOP models derived from signatures of oncogenic pathways associated with activation of BMI1; Myc; Her2/neu; Ras; beta-catenin; Suz12; E2F; and CCND1 oncogenes. Application of a CTOP algorithm to large databases of early-stage breast and prostate tumors identifies cancer patients with 100% probability of a cure with existing cancer therapies as well as patients with nearly 100% likelihood of treatment failure, thus providing a clinically feasible framework essential for introduction of rational evidence-based individualized therapy selection and prescription protocols. Our analysis indicates that genetic determinants of human disease susceptibility and severity are encoded by population differentiation SNP variants. Evolution of these SNPs is driven by geographically-localized form of natural selection causing population differentiation. Recent analysis identifies a class of SNPs regulating gene expression in normal individuals and likely determining unique genome-wide expression profiles of each individual. We propose that critical disease-causing combinations of SNP variants arise from SNPs regulating mRNA levels and determining genome-wide haplotype patterns of individual's disease susceptibility.

Biomarkers, Tumor↗

Accuracy of haplotype frequency estimation for biallelic loci, via the expectation-maximization algorithm for unphased diploid genotype data.

Haplotype analyses have become increasingly common in genetic studies of human disease because of their ability to identify unique chromosomal segments likely to harbor disease-predisposing genes. The study of haplotypes is also used to investigate many population processes, such as migration and immigration rates, linkage-disequilibrium strength, and the relatedness of populations. Unfortunately, many haplotype-analysis methods require phase information that can be difficult to obtain from samples of nonhaploid species. There are, however, strategies for estimating haplotype frequencies from unphased diploid genotype data collected on a sample of individuals that make use of the expectation-maximization (EM) algorithm to overcome the missing phase information. The accuracy of such strategies, compared with other phase-determination methods, must be assessed before their use can be advocated. In this study, we consider and explore sources of error between EM-derived haplotype frequency estimates and their population parameters, noting that much of this error is due to sampling error, which is inherent in all studies, even when phase can be determined. In light of this, we focus on the additional error between haplotype frequencies within a sample data set and EM-derived haplotype frequency estimates incurred by the estimation procedure. We assess the accuracy of haplotype frequency estimation as a function of a number of factors, including sample size, number of loci studied, allele frequencies, and locus-specific allelic departures from Hardy-Weinberg and linkage equilibrium. We point out the relative impacts of sampling error and estimation error, calling attention to the pronounced accuracy of EM estimates once sampling error has been accounted for. We also suggest that many factors that may influence accuracy can be assessed empirically within a data set-a fact that can be used to create "diagnostics" that a user can turn to for assessing potential inaccuracies in estimation.

Algorithms↗

[Association between methylenetetrahydrofolate reductase gene haplotypes and the susceptibility of chromosomal damage in coke-oven workers].

OBJECTIVE: To investigate the association between MTHFR gene variances and chromosomal damage levels in peripheral blood lymphocyte in coke-oven workers exposed to polycyclic aromatic hydrocarbons (PAHs). METHODS: One-hundred and forty coke-oven workers who exposed to a high level of PAHs and sixty-six non-exposed controls were selected as the study subjects. Chromosomal damage in peripheral lymphocyte was measured by the cytokinesis-block micronucleus (CBMN) assay. Urinary 1-hydroxypyrene (1-OHP) levels were measured as the internal dose of PAHs exposure. Two single nucleotide polymorphisms(SNPs) in MTHFR gene, including C677T, A1298C were detected by PCR-RFLP. The MTHFR haplotypes were estimated by Bayesian statistical method with the software of PHASE Version 2.1. The associations between haplotype pairs and CBMN were assessed by analysis of covariance in the coke-oven workers and controls. RESULTS: The variant allele frequencies for MTHFRC677T and A1298C were 0.56 and 0.16 respectively, which consistent with Hardy-Weinberg equilibrium. There was linkage disequilibrium between the two SNPs (D' = 0.99) in this study. Four haplotypes were calculated by PHASE, in terms of 677T - 1298A, 677C-1298A, 677C-1298C and 677T-1298C, the frequencies were 0.555,0.279,0.163 and 0.003 respectively. In coke-oven workers, the frequencies of total micronucleus of non-677C-1298A/677C-1298A haplotype pair was significantly higher than 677C-1298A/677C-1298A (1.00 +/- 0.67 vs 0.60 +/- 0.41, P = 0.04). The frequencies of total micronucleus of 677T-1298A/677T-1298A haplotype pair was significantly higher than 677C-1298A/677C-1298A (1.08 +/- 0.71 vs 0.60 +/- 0.41, P = 0.04). In coke-oven workers, the frequencies of total micronucleus among the different SNPs were not significant differences, either in the controls. CONCLUSION: The haplotypes of MTHFR gene might be one genetic susceptibility factors of PAH induced chromosomal damage in coke-oven workers.

Air Pollutants, Occupational↗

Inference and analysis of haplotypes from combined genotyping studies deposited in dbSNP.

In the attempt to understand human variation and the genetic basis of complex disease, a tremendous number of single nucleotide polymorphisms (SNPs) have been discovered and deposited into NCBI's dbSNP public database. More than 2.7 million SNPs in the database have genotype information. This data provides an invaluable resource for understanding the structure of human variation and the design of genetic association studies. The genotypes deposited to dbSNP are unphased, and thus, the haplotype information is unknown. We applied the phasing method HAP to obtain the haplotype information, block partitions, and tag SNPs for all publicly available genotype data and deposited this information into the dbSNP database. We also deposited the orthologous chimpanzee reference sequence for each predicted haplotype block computed using the UCSC BLASTZ alignments of human and chimpanzee. Using dbSNP, researchers can now easily perform analyses using multiple genotype data sets from the same genomic regions. Dense and sparse genotype data sets from the same region were combined to show that the number of common haplotypes is significantly underestimated in whole genome data sets, while the predicted haplotypes over the common SNPs are consistent between studies. To validate the accuracy of the predictions, we bench-marked HAP's running time and phasing accuracy against PHASE. Although HAP is slightly less accurate than PHASE, HAP is over 1000 times faster than PHASE, making it suitable for application to the entire set of genotypes in dbSNP.

Animals↗

Haplotype reconstruction and estimation of haplotype frequencies from nuclear families with only one parent available.

Recent literature has suggested that haplotype inference through close relatives, especially from nuclear families can be an alternative strategy in determining the linkage phase. In this paper, haplotype reconstruction and estimation of haplotype frequencies via expectation maximization (EM) algorithm including nuclear families with only one parent available is proposed. Parent and his (her) child are treated as parent-child pair with one shared haplotype. This reduces the number of potential haplotype pairs for both parent and child separately, resulting in a higher accuracy of the estimation. In a series of simulations, the comparisons of PHASE, GENEHUNTER, EM-based approach for complete nuclear families and our approach are carried out. In all situations, EM-based approach for trio data is comparable but slightly worse error rate than PHASE, our approach is slightly better and much faster than PHASE for incomplete trios, the performance of GENEHUNTER is very bad in simple nuclear family settings and dramatically decreased with the number of markers being increased. On the other hand, the comparison result of different sampling designs demonstrates that sampling trios is the most efficient design to estimate haplotype frequencies in populations under same genotyping cost.

Child↗

Full haplotype-mismatched hematopoietic stem-cell transplantation: a phase II study in patients with acute leukemia at high risk of relapse.

PURPOSE: Establishment of hematopoietic stem-cell (HSC) transplantation from mismatched relatives is feasible for patients with acute leukemia. As our original method of graft processing was unsuitable for large-scale clinical studies, we use automated devices for CD34+ cell purification. PATIENTS AND METHODS: Sixty-seven patients with acute myeloid leukemia (AML; 19 complete remission [CR] 1, 14 CR 2, nine CR > 2, 25 in relapse) and 37 with acute lymphoid leukemia (ALL; 14 CR 1, eight CR 2, two CR > 2, 13 in relapse) were conditioned with total-body irradiation, thiotepa, fludarabine, and antithymocyte globulin. Peripheral-blood progenitor cells were mobilized with recombinant human granulocyte colony-stimulating factor and depleted of T-cells using CD34+ cell immunoselection. No post-transplantation graft-versus-host disease (GvHD) prophylaxis was administered. RESULTS: Primary engraftment was achieved in 94 of 101 assessable patients. Six of the seven patients who rejected the primary graft, engrafted after a second transplantation. Overall, 100 of 101 patients engrafted. Acute GvHD developed in eight of 100 patients, and chronic GvHD, in five of 70 assessable patients. Thirty-eight patients died of nonleukemic causes. Relapse occurred in nine of 66 patients receiving transplantation in remission and in 17 of 38 receiving transplantation in relapse. Median follow-up of the 40 patients who survived event-free was 22 months (range, 1 to 65 months). Event-free survival (+/- standard deviation) rate was 48% +/- 8% and 46% +/- 10%, respectively, for the 42 AML and 24 ALL patients receiving transplantation in remission. CONCLUSION: Our transplantation procedure provides reliable, reproducible CD34+ cell purification, high engraftment rates, and prevention of GvHD. The mismatched-related transplant emerges as a viable, alternative source of stem cells for acute leukemia patients without matched donors and/or those who urgently need transplantation.

Adolescent↗

SNPs and snails and puppy dogs' tails: analysis of SNP haplotype data using the gamete competition model.

The gamete competition model is a likelihood version of the transmission disequilibrium test (TDT) that is inspired by conditional logistic regression and the Bradley-Terry ranking procedure. In family-based association studies, both the TDT and the gamete competition model apply directly to data on a single nucleotide polymorphism (SNP). Because any given SNP has limited polymorphism, it is tempting to collect several SNPs within a gene into a single super marker whose alleles are haplotypes. Unfortunately, this tactic wreaks havoc with the traditional TDT, which requires codominant markers (Spielman et al. 1993; Terwilliger & Ott, 1992). Eliminating phase ambiguities by assigning haplotypes to individuals before conducting the TDT may give misleading results because only the most probable haplotypes are then considered. Because pedigree implementations of the gamete competition model can accommodate dominant as well as codominant markers, they circumvent the phase problem by including all possible phases weighted by their estimated frequencies.

Germ Cells↗

GERBIL: Genotype resolution and block identification using likelihood.

The abundance of genotype data generated by individual and international efforts carries the promise of revolutionizing disease studies and the association of phenotypes with individual polymorphisms. A key challenge is providing an accurate resolution (phasing) of the genotypes into haplotypes. We present here results on a method for genotype phasing in the presence of recombination. Our analysis is based on a stochastic model for recombination-poor regions ("blocks"), in which haplotypes are generated from a small number of core haplotypes, allowing for mutations, rare recombinations, and errors. We formulate genotype resolution and block partitioning as a maximum-likelihood problem and solve it by an expectation-maximization algorithm. The algorithm was implemented in a software package called GERBIL (genotype resolution and block identification using likelihood), which is efficient and simple to use. We tested GERBIL on four large-scale sets of genotypes. It outperformed two state-of-the-art phasing algorithms. The phase algorithm was slightly more accurate than GERBIL when allowed to run with default parameters, but required two orders of magnitude more time. When using comparable running times, GERBIL was consistently more accurate. For data sets with hundreds of genotypes, the time required by phase becomes prohibitive. We conclude that GERBIL has a clear advantage for studies that include many hundreds of genotypes and, in particular, for large-scale disease studies.

Chromosome Mapping↗

Power of direct vs. indirect haplotyping in association studies.

Haplotype analysis is essential to studies of the genetic factors underlying human disease, but requires a large sample size of phase-known data. Recently, directly haplotyping individuals was suggested as a means of maximizing the phase-known data from a sample. Haplotyping, however, is much more labor-intensive than indirectly inferring haplotypes from genotypes (genotyping). This study uses simulations to compare the power of each methodology to detect associations between a haplotype and a trait or disease locus under conditions of varying linkage disequilibrium. The relative power of haplotyping over genotyping in association studies increases with decreasing sample size, decreasing linkage disequilibrium, increasing [corrected] numbers of marker loci, and decreasing numbers of different haplotypes. In addition, the frequency of the haplotype of interest and the magnitude of its association with the disease affect the power. From a cost-benefit standpoint, genotyping would be favored with large multiplicative risks (relative risk of haplotype >2.5). If case numbers are limiting rather than cost, haplotyping would maximize the information obtained. At small haplotype frequencies (e.g., <0.05), haplotyping is relatively more efficient, but there is little absolute power to detect associations under either methodology. Given the much larger laboratory resources required for direct haplotyping, genotyping would probably be favored under most conditions, but this must be balanced against the unit costs associated with recruitment and phenotyping. In the context of multipurpose, prospective cohort studies (e.g., the UK Biobank study), there may be a general value in establishing a series of directly haplotyped individuals to serve as controls for a number of alternative studies.

Algorithms↗

Molecular haplotyping of genetic markers 10 kb apart by allele-specific long-range PCR.

Haplotypes, combinations of polymorphic markers in a chromosome, are critical for genome diversity research. However, their utility in population samplings is compromised by uncertain linkage phase determinations from unrelated individuals. Molecular haplotyping accomplishes direct phase determination by generation of hemizygous templates from diploid genomic samples. We report molecular haplotyping by allele-specific long-range PCR of two markers 9.5 kb apart at the CD4 locus: a bi-allelic Alu deletion and a multi-allelic repeat. We verified CD4 molecular haplotypes by classical Mendelian analysis. Molecular haplotyping should prove useful in mapping disease genes and in establishing founder effects.

Alleles↗

Molecular haplotyping at high throughput.

Reconstruction of haplotypes, or the allelic phase, of single nucleotide polymorphisms (SNPs) is a key component of studies aimed at the identification and dissection of genetic factors involved in complex genetic traits. In humans, this often involves investigation of SNPs in case/control or other cohorts in which the haplotypes can only be partially inferred from genotypes by statistical approaches with resulting loss of power. Moreover, alternative statistical methodologies can lead to different evaluations of the most probable haplotypes present, and different haplotype frequency estimates when data are ambiguous. Given the cost and complexity of SNP studies, a robust and easy-to-use molecular technique that allows haplotypes to be determined directly from individual DNA samples would have wide applicability. Here, we present a reliable, automated and high-throughput method for molecular haplotyping in 2 kb, and potentially longer, sequence segments that is based on the physical determination of the phase of SNP alleles on either of the individual paternal haploids. We demonstrate that molecular haplotyping with this technique is not more complicated than SNP genotyping when implemented by matrix-assisted laser desorption/ionisation mass spectrometry, and we also show that the method can be applied using other DNA variation detection platforms. Molecular haplotyping is illustrated on the well-described beta(2)-adrenergic receptor gene.

Alleles↗