Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Main haplotypes and mutational analysis of vitamin K epoxide reductase (VKORC1) in a Swedish population: a retrospective analysis of case records.

BACKGROUND: Vitamin K epoxide reductase (VKORC1) is the site of inhibition by coumarins. Several reports have shown that mutations in the gene encoding VKORC1 affect the sensitivity of the enzyme for warfarin. Recently, three main haplotypes of VKORC1; *2, *3 and *4 have been observed, that explain most of the genetic variability in warfarin dose among Caucasians. OBJECTIVES: We have investigated the main haplotypes of the VKORC1 gene in a Swedish population. Additional objective was to screen the studied population for mutations in the coding region of VKORC1 gene. PATIENTS/METHODS: Warfarin doses and plasma S- and R-warfarin of 98 patients [with a target International Normalized Ratio (INR) of 2.0-3.0] have been correlated to VKORC1 haplotypes. Controls of 180 healthy individuals have also been haplotyped. Furthermore, a retrospective analysis of case records was performed to find any evidence indicating influence of VKORC1 haplotypes on warfarin response in the first 4 weeks (initiation phase) and the latest 12 months of warfarin treatment. RESULTS AND CONCLUSIONS: Our result shows that VKORC1*2 is the most important haplotype for warfarin dosage. Patients with VKORC1*2 haplotype had more frequent visits than patients with VKORC1*3 or *4 haplotypes, higher coefficient of variation (CV) of prothrombin time-INR and higher percentage of INR values outside the therapeutic interval (i.e. 2.0-3.0) than patients with VKORC1*3 or *4 haplotypes. Also, there was a statistically significant difference in warfarin dose (P < 0.001) and R-warfarin plasma levels (P < 0.01) between VKORC1*2 and VKORC1*3 or 4 haplotypes. Patients with VKORC1*2 haplotype seem to require much lower warfarin doses than other patients.

Adult↗

Inference of haplotypes from samples of diploid populations: complexity and algorithms.

The next phase of human genomics will involve large-scale screens of populations for significant DNA polymorphisms, notably single nucleotide polymorphisms (SNPs). Dense human SNP maps are currently under construction. However, the utility of those maps and screens will be limited by the fact that humans are diploid and it is presently difficult to get separate data on the two "copies." Hence, genotype (blended) SNP data will be collected, and the desired haplotype (partitioned) data must then be (partially) inferred. A particular nondeterministic inference algorithm was proposed and studied by Clark (1990) and extensively used by Clark et al. (1998). In this paper, we more closely examine that inference method and the question of whether we can obtain an efficient, deterministic variant to optimize the obtained inferences. We show that the problem is NP-hard and, in fact, Max-SNP complete; that the reduction creates problem instances conforming to a severe restriction believed to hold in real data (Clark, 1990); and that even if we first use a natural exponential-time operation, the remaining optimization problem is NP-hard. However, we also develop, implement, and test an approach based on that operation and (integer) linear programming. The approach works quickly and correctly on simulated data.

Algorithms↗

Old and new genetics help ordering loci at the telomere of the human X-chromosome long arm.

A Sardinian pedigree described in 1964 for having been found to segregate at the X-linked loci for the Xga antigen, G6PD deficiency, Protan and Deutan color blindness, with an instance of recombination between the last two loci, was re-examined with respect to four common X-linked DNA polymorphisms detected by molecular probes homologous to critical subregions of the human X chromosome. Two branches of this pedigree--including the one with the Protan-Deutan recombinant--were found to segregate also for the common BamHI polymorphism identified with the cDNA probe pHPT-2 or the HPRT gene (Xq26). The analysis of the chromosome haplotypes in the male offspring of the phase known penta-heterozygous mother suggests that the probable order of the relevant loci is HPRT, Deutan, G6PD, Protan, Xq telomere. Though we are fully aware of the risks of generalizing the significance of observations made on a single exceptional pedigree, we believe that this report outlines the potential of families of the type described as research tools to resolve the linear order of tightly X-linked loci and to investigate the biology of genetic recombination in humans.

Blood Group Antigens↗

Efficient cytoplasmic delivery of a fluorescent dye by pH-sensitive immunoliposomes.

We previously showed that liposomes composed of dioleoylphosphatidyl-ethanolamine and palmitoyl-homocysteine (8:2) are highly fusion competent when exposed to an acidic environment of pH less than 6.5. (Connor, J., M. B. Yatvin, and L. Huang, 1984, Proc. Natl. Acad. Sci. USA. 81:1715-1718). Palmitoyl anti-H2Kk was incorporated into these pH-sensitive liposomes by a modified reserve-phase evaporation method. Mouse L929 cells (k haplotype) treated with immunoliposomes composed of dioleoylphosphatidylethanolamine/palmitoyl-homocysteine (8:2) with an entrapped fluorescent dye, calcein, showed diffused fluorescence throughout the cytoplasm. Measurements by use of a microscope-associated photometer gave an approximate value of 50 microM for the cytoplasmic calcein concentration. This concentration represents an efficient delivery of the aqueous content of the immunoliposome. Cells treated with immunoliposomes composed of dioleoylphosphatidylcholine (pH-insensitive liposomes) showed only punctate fluorescence. The cytoplasmic delivery of calcein by the pH-sensitive immunoliposomes could be inhibited by chloroquine or by incubation at 20 degrees C. These results suggest that the efficient cytoplasmic delivery involves the endocytic pathway, particularly the acidic organelles such as the endosomes and/or lysosomes. One possibility is that the immunoliposomes fuse with the endosome membranes from within the endosomes, thus releasing the contents into the cytoplasm. This nontoxic method should be widely applicable to the intracellular delivery of biomolecules into living cells.

Animals↗

A haplotype-resolved pangenome of the barley wild relative Hordeum bulbosum.

Wild plants can contribute valuable genes to their domesticated relatives1. Fertility barriers and a lack of genomic resources have hindered the effective use of crop-wild introgressions. Decades of research into barley's closest wild relative, Hordeum bulbosum, a grass native to the Mediterranean basin and Western Asia, have yet to manifest themselves in the release of a cultivar bearing alien genes2. Here we construct a pangenome of bulbous barley comprising 10 phased genome sequence assemblies amounting to 32 distinct haplotypes. Autotetraploid cytotypes, among which the donors of resistance-conferring introgressions are found, arose at least twice, and are connected among each other and to diploid forms through gene flow. The differential amplification of transposable elements after barley and H.&#x2009;bulbosum diverged from each other is responsible for genome size differences between them. We illustrate the translational value of our resource by mapping non-host resistance to a viral pathogen to a structurally diverse multigene cluster that has been implicated in diverse immune responses in wheat and barley.

Hordeum↗

Fine genetic mapping using haplotype analysis and the missing data problem.

The genetic basis of many human diseases, especially those with substantial genetic determinants, has been identified. Notable amongst others are cystic fibrosis, Huntington's disease and some forms of cancer. However, the detection of genetic factors with more modest effects such as in bipolar disorders and a majority of the cancers, has been more complicated. Standard linkage analysis procedures may not only have little power to detect such genes but they do, at best, only narrow the location of the disease susceptibility gene to a rather large region. Association studies are therefore necessary to further unveil the aetiological relevance of these factors to disease. However, the number of tests required if such procedures were used in extended genome-wide screens, is prohibitive and as such association studies have seen limited application, except in the investigation of candidate genes. In this paper, we discuss a logistic regression approach as a generalization of this procedure so that it can accommodate clusters of linked markers or candidate genes. Furthermore, we introduce an expectation maximization (E-M) algorithm with which to estimate haplotype frequencies for multiple locus systems with incomplete information on phase.

Algorithms↗

Risk of small-for-gestational age is associated with common anti-inflammatory cytokine polymorphisms.

BACKGROUND: Anti-inflammatory cytokines play a key role in pregnancy maintenance. Genetic variation in anti-inflammatory cytokines could influence a woman's risk of adverse reproductive outcomes. METHODS: We investigated the relationship of polymorphisms in interleukin 4 (IL4), IL5, IL10, IL13, and transforming growth factor (TGFbeta1) with spontaneous preterm birth and small-for-gestational age (SGA) in a nested case-control study of a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. Semi-Bayesian hierarchical logistic regression was used to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: African-American mothers who carried the IL4 GCC haplotype had greater risk of spontaneous preterm birth (OR = 2.9; 95% CI = 1.2-7.4). In white mothers, carriers of the "low-producing" IL4 CC and IL10 ATA haplotypes had markedly reduced risk of SGA (for the CC haplotype, 0.2 [0.0-1.2]; for the ATA haplotype, 0.5 [0.3-0.8]), whereas carriers of the "high-producing" IL4(-589)T variant had increased risk of SGA in both African-American and white mothers. CONCLUSIONS: Variants related to decreased anti-inflammatory cytokine production may lower risk of SGA. Furthermore, the same mechanism that protects against SGA might increase risk of spontaneous preterm birth.

Black or African American↗

Population differences in haplotype structure within a human olfactory receptor gene cluster.

We investigated the population differences in patterns of single nucleotide polymorphisms (SNPs) for a 400 kb olfactory receptor (OR) gene cluster on human chromosome 17p13.3. Samples were drawn from 35 individuals, of four different ethnogeographical origins: Pygmies, Bedouins, Yemenite Jews and Ashkenazi Jews. Of the 74 SNPs identified, two segregated between pseudogenized and intact ORs, while a third involved a change in a highly conserved motif proposed to mediate ligand-induced signal transduction. Linkage disequilibrium (LD) was computed based on phase inference across the cluster using Clark's haplotype subtraction algorithm. We also calculated LD directly from the genotypes using the expectation-maximization (EM) algorithm. Both methods yielded very similar results. Our analyses revealed substantial differences in nucleotide diversity, haplotype distribution and LD patterns among the different human populations. In particular, the two Jewish populations had low haplotype diversity and negligible decay of LD across the entire genomic region. Intriguingly, the three functional SNPs segregated at different frequencies in the different ethnogeographical groups, with the Pygmies having higher frequencies of the intact OR genes. Our data suggests that OR genes may have evolved to create different functional repertoires in distinct human populations.

Chromosomes, Human, Pair 17↗

Common genetic polymorphisms in the 5'-flanking region of the SULT1A1 gene: haplotypes and their association with platelet enzymatic activity.

SULT1A1 is a phase II detoxification enzyme involved in the biotransformation of a wide variety of endogenous and exogenous phenolic compounds. Human platelet SULT1A1 enzymatic activity shows marked inter-individual variability and a common coding polymorphism, SULT1A1*1/*2, has been described that accounts for a proportion of this variability. We examined the 5'-flanking region of the SULT1A1 gene to determine if genetic variability in this portion of the gene influenced enzymatic activity. Direct sequencing revealed five common genetic polymorphisms (-624G>C, -396G>A, -358A>C, -341C>G and -294T>C) that were present at different allele frequencies in Caucasian, African-American and Chinese groups. Platelet SULT1A1 enzymatic activity was significantly correlated with individual promoter region polymorphisms and the associations were different between African-Americans and Caucasians. Haplotypes were constructed and platelet enzymatic activity according to haplotype was examined. The haplotypes were also significantly correlated with activity; haplotypes GAACT and GGACT (accounting for 13% and 5% of inter-individual variability in platelet activity, respectively) were important in Caucasians while haplotypes GAACC, GAACT and GGACC (accounting for 8%, 5% and 4% of variability) were significantly associated with activity in African-Americans. The coding region polymorphism, SULT1A1*1/*2 was in linkage disequilibrium with the promoter region polymorphisms and showed no effect on activity when examined in the context of the 5'-flanking region polymorphisms. These studies indicate that variation in the promoter region of the SULT1A1 gene exerts a significant influence on enzymatic activity.

5' Flanking Region↗

Adjacent genes, for COL2A1 and the vitamin D receptor, are associated with separate features of radiographic osteoarthritis of the knee.

OBJECTIVE: To study the association of the COL2A1 genotype, in relation to the vitamin D receptor (VDR) genotype, with features of radiographic osteoarthritis (ROA) in a population of elderly men and women. METHODS: In this cross-sectional study, we analyzed a population-based sample of 851 men and women ages 55-80 years from a large cohort study, the Rotterdam Study. We determined the prevalence of ROA of the knee according to the Kellgren/Lawrence (K/L) score and features of ROA (presence of osteophytes and narrowing of the joint space [JSN]) without considering clinical parameters of the disease. Genotypes were determined at a variable-number tandem repeats marker 1 kb downstream of the COL2A1 gene using a newly developed heteroduplexing method. The VDR genotype was previously determined by a direct molecular haplotyping polymerase chain reaction method to establish the phase of alleles at 3 adjacent restriction fragment length polymorphisms for Bsm I, Apa I, and Taq I. RESULTS: We found the COL2A1 genotype to be associated with a 2-fold increased risk for JSN, but not with osteophytes or the K/L score. We had previously found the VDR genotype to be associated with osteophytes and the K/L score, but not with JSN. When the COL2A1 genotype was analyzed in combination with the VDR genotype, we found evidence suggesting that the presence of haplotypes of the 2 genes was associated with increased risk for ROA. CONCLUSION: Our findings demonstrate that both the COL2A1 gene and the VDR gene are involved in ROA, but in separate features. The COL2A1 genotype is associated with JSN, while the VDR genotype is associated with osteophytes.

Aged↗

Vascular endothelial growth factor gene polymorphisms and risk of primary lung cancer.

Angiogenesis is an essential process in the development, growth, and metastasis of malignant tumors including lung cancer. DNA sequence variations in the vascular endothelial growth factor (VEGF) gene may lead to altered VEGF production and/or activity, thereby causing interindividual differences in the susceptibility to lung cancer via their actions on the pathways of tumor angiogenesis. To test this hypothesis, we investigated the potential association between three VEGF polymorphisms (-460T > C, +405C > G, and 936C > T)/haplotypes and the risk of lung cancer in a Korean population. VEGF genotypes were determined in 432 lung cancer patients and 432 healthy controls that were frequency matched for age and sex. VEGF haplotypes were predicted using Bayesian algorithm in the phase program. Compared with the combined +405 CC and CG genotype, the +405 GG genotype found associated with a significantly decreased risk of small cell carcinoma [SCC; adjusted odds ratio (OR), 0.36; 95% confidence interval (95% CI), 0.17-0.78]. The 936 CT genotype and the combined 936 CT and TT genotype were also associated with a significantly decreased risk of SCC compared with the 936 CC genotype (adjusted OR, 0.47; 95% CI, 0.26-0.85 and adjusted OR, 0.44; 95% CI, 0.24-0.80, respectively). Haplotype CGT was associated with a significantly decreased risk of SCC (adjusted OR, 0.39; 95% CI, 0.18-0.87), whereas haplotype TCC conferred a significantly increased risk of SCC (adjusted OR, 1.63; 95% CI, 1.14-2.33). None of the VEGF polymorphisms studied significantly influenced the susceptibility to lung cancer except SCC. However, haplotypes TCT and TGT were significantly associated with the risk of overall lung cancer, respectively (adjusted OR, 0.38; 95% CI, 0.25-0.60 and adjusted OR, 3.94; 95% CI, 2.00-7.76, respectively). These effects of haplotypes TCT and TGT on lung cancer risk were observed in three major histologic types of lung cancer. These results suggest that the VEGF gene may be contribute to an inherited predisposition to lung cancer.

Aged↗

Effect of itraconazole on the pharmacokinetics and pharmacodynamics of fexofenadine in relation to the MDR1 genetic polymorphism.

OBJECTIVE: Our objective was to evaluate the effect of itraconazole, a P-glycoprotein inhibitor, on the pharmacokinetics and pharmacodynamics of fexofenadine, a P-glycoprotein substrate, in relation to the multidrug resistance 1 gene (MDR1) G2677T/C3435T haplotype. METHODS: A single oral dose of 180 mg fexofenadine was administered to 7 healthy subjects with the 2677GG/3435CC (G/C) haplotype and 7 with the 2677TT/3435TT (T/T) haplotype. One hour before the fexofenadine dose, either 200 mg itraconazole or placebo was administered to the subjects in a double-blinded, randomized, crossover manner with a 2-week washout period. Histamine-induced wheal and flare reactions were measured to assess the effects on the antihistamine response. RESULTS: In the placebo phase, pharmacokinetic parameters of fexofenadine showed no statistically significant difference between 2 MDR1 haplotypes; the area under the curve from time 0 to infinity (AUC(0-infinity)) of fexofenadine in the T/T and G/C groups was 5194.0 +/- 1910.8 and 4040.4 +/- 1832.2 ng.mL(-1).h(-1), respectively (P = .271), and the oral clearance (CL/F) was 530.9 +/- 191.1 and 806.0 +/- 355.3 mL.h(-1).kg(-1), respectively (P = .096). The disposition of itraconazole, a substrate of P-glycoprotein, was not significantly different between the 2 haplotypes. After itraconazole pretreatment, however, the differences in fexofenadine pharmacokinetics became statistically significant; the mean fexofenadine AUC(0-infinity) in the T/T group was significantly higher than that in the G/C group (15,630.6 +/- 5070.0 and 9252.9 +/- 2044.1 ng/mL.h, respectively; P = .007), and CL/F of the T/T subjects was lower than that of the G/C subjects (167.0 +/- 33.3 and 292.3 +/- 42.2 mL.h(-1).kg(-1), respectively; P < .001). Itraconazole pretreatment caused more than a 3-fold increase in the peak concentration of fexofenadine and the area under the curve to 6 hours compared with the placebo phase. This resulted in a significantly higher suppression of the histamine-induced wheal and flare reactions in the itraconazole pretreatment phase compared with those in the placebo phase. CONCLUSION: The effect of MDR1 G2677T/C3435T haplotypes on fexofenadine disposition are magnified in the presence of itraconazole. Itraconazole pretreatment significantly altered the disposition of fexofenadine and thus its peripheral antihistamine effects.

Adult↗

Pedigree disequilibrium tests for multilocus haplotypes.

Association tests of multilocus haplotypes are of interest both in linkage disequilibrium mapping and in candidate gene studies. For case-parent trios, I discuss the extension of existing multilocus methods to include ambiguous haplotypes in tests of models which distinguish between the cis and trans phase. A likelihood-ratio test is proposed, using the expectation-maximization (E-M) algorithm to account for haplotype ambiguities. Assumptions about the population structure are required, but realistic situations, including population stratification, which violate the assumptions lead to conservative tests. I describe a permutation procedure for the null hypothesis of interest, which controls for violation of the assumptions. For general pedigrees, I describe extensions of the pedigree disequilibrium test to include uncertain haplotypes. The summary statistics are replaced by their expected values over prior distributions of haplotype frequencies. If prior distributions are not available, a valid test is possible by using the E-M algorithm to estimate the null distribution of haplotype frequencies. Similar methods are available for quantitative traits. Exact permutation tests are difficult to construct in small samples, but an approximate procedure is appropriate in large samples, and can be used to account for dependencies between tests of multiple haplotypes and loci.

Algorithms↗

Detecting disease associations due to linkage disequilibrium using haplotype tags: a class of tests and the determinants of statistical power.

In the 'indirect' method of detecting genetic associations between a trait and a DNA variant, we type several markers in a gene or chromosome region of linkage disequilibrium. If there is association between markers and the trait, we presume the existence of one or more causal polymorphisms in the region. In order to obtain a sufficiently dense set of markers it will almost always be necessary to use single nucleotide polymorphisms (SNPs). Although there is an emerging literature on methods for choosing an optimal set of 'haplotype tag SNPs' (htSNPs) to detect association between a genetic region and a trait, less attention has been given to the problem of how such studies should be analysed when completed, and how the initial data which was used to select the htSNPs should be incorporated into the analysis. This paper discusses this problem for both population- and family-based association studies. The role of the R2 measure of association between a causal locus and various methods of scoring of marker haplotypes is highlighted. In most cases, the simplest method of scoring (locus coding), which does not require phase resolution, is shown generally to be more powerful than scoring methods that include haplotype information. A new 'multi-locus TDT' is also proposed.

Data Interpretation, Statistical↗

Polymorphisms in TGF-beta1 gene and the risk of lung cancer.

BACKGROUND: Transforming growth factor-beta1 (TGF-beta1) functions as a suppressor of tumor initiation by inhibiting cellular proliferation or by promoting cellular differentiation or apoptosis in the early phase of cancer development. Variations in the DNA sequence in the TGF-beta1 gene may lead to altered TGF-beta1 production and/or activity, and so this can modulate an individual's susceptibility to lung cancer. To test this hypothesis, we investigated the association of the TGF-beta1 -509C > T and 869T > C (L10P) polymorphisms and their haplotypes with the risk of lung cancer in a Korean population. METHODS: The TGF-beta1 genotypes were determined in 432 lung cancer patients and in 432 healthy control subjects who were frequency-matched for age and gender. The TGF-beta1 haplotypes were predicted using a Bayesian algorithm in the Phase program. RESULTS: Individuals with at least one -509T allele were at a significantly decreased risk of adenocarcinoma (AC) and small cell carcinoma (SM), as compared with carriers with the -509CC genotype [adjusted odds ratio (OR), 0.63; 95% confidence interval (CI), 0.42-0.96; P = 0.04; and adjusted OR, 0.45; 95% CI, 0.27-0.76; P = 0.002; respectively]. For the 869T > C polymorphism, the combined TC + CC genotype was associated with a significantly decreased risk of SM compared with the TT genotype (adjusted OR, 0.52; 95% CI, 0.31-0.88; P = 0.01). Consistent with the results of the genotyping analyses, the -509T/869C haplotype was associated with a significantly decreased risk of AC and SM as compared with the -509C/869T haplotype (adjusted OR, 0.75; 95% CI, 0.57-0.98; P = 0.04; and adjusted OR, 0.67; 95% CI, 0.47-0.96; P = 0.02; respectively). CONCLUSION: The TGF-beta1 -509C > T and 869T > C polymorphisms and their haplotypes may contribute to genetic susceptibility to AC and SM of the lung.

Adenocarcinoma↗

Relative efficiency of ambiguous vs. directly measured haplotype frequencies.

Haplotypes are useful for both fine-mapping of susceptibility loci and evaluation of sequence variation at multiple sites along a chromosome. However, they are difficult to directly measure over long stretches of DNA in diploid organisms. Consequently, multiple genetic markers are typically measured, without linkage phase information, giving rise to a subject's diplotype. From diplotype data, haplotypes are often inferred by pedigree information, or treated as partially missing data when haplotype frequencies are estimated among unrelated subjects. This latter ambiguity can increase the variance of the estimated haplotype frequencies. Douglas et al. ([2001] Nat. Genet. 28:361-364) recently quantified the relative efficiency of estimating haplotype frequencies from the diplotypes of unrelated subjects, relative to directly measured haplotypes via somatic cell hybrids (conversion technology), and demonstrated that unknown linkage phase can lead to a large loss of efficiency. However, their results were based on linkage equilibrium among marker loci, which may not be realistic for closely linked markers. We extend their relative efficiency calculations by several aspects: 1) allowance for linkage disequilbrium (LD) among marker loci; 2) evaluation of different patterns of LD; and 3) evaluation of nuclear families with and without parents. We show that although the loss in efficiency of haplotype frequencies among unrelated subjects decreases as LD increases to its maximum value, the general conclusions of Douglas et al. ([2001] Nat. Genet. 28:361-364) hold true for a variety of LD patterns and magnitudes. However, our results also demonstrate that trios of parents+one child are highly efficient for haplotype frequency estimation, that additional children offer little information, and that siblings without parents can be grossly inefficient. Genet. Epidemiol. 23:426-443, 2002.

Algorithms↗

Genome-wide definitive haplotypes determined using a collection of complete hydatidiform moles.

We present genome-wide definitive haplotypes, determined using a collection of 74 Japanese complete hydatidiform moles, each carrying a genome derived from a single sperm. The haplotypes incorporate 281,439 common SNPs, genotyped with a high throughput array-based oligonucleotide hybridization technique. Comparison of haplotypes inferred from pseudoindividuals (constructed from randomized mole pairs) with those of moles showed some switch errors in resolution of phases by the computational inference method. The effects of these errors on local haplotype structure and selection of tag SNPs are discussed. We also show that definitive haplotypes of moles may be useful for elucidation of long-range haplotype structure, and should be more effective for detecting extended haplotype homozygosity indicative of positive selection.

Female↗

Statistical estimation and pedigree analysis of CCR2-CCR5 haplotypes.

As more SNP marker data becomes available, researchers have used haplotypes of markers, rather than individual polymorphisms, for association analysis of candidate genes. In order to perform haplotype analysis in a population-based case-control study, haplotypes must be determined by estimation in the absence of family information or laboratory methods for establishing phase. Here, we test the accuracy of the Expectation-Maximization (EM) algorithm for estimating haplotype state and frequency in the CCR2-CCR5 gene region by comparison with haplotype state and frequency determined by pedigree analysis. To do this, we have characterized haplotypes comprising alleles at seven biallelic loci in the CCR2-CCR5 chemokine receptor gene region, a span of 20 kb on chromosome 3p21. Three-generation CEPH families (n=40), totaling 489 individuals, were genotyped by the 5'nuclease assay (TaqMan). Haplotype states and frequencies were compared in 103 grandparents who were assumed to have mated at random. Both pedigree analysis and the EM algorithm yielded the same small number of haplotypes for which linkage disequilibrium was nearly maximal. The haplotype frequencies generated by the two methods were nearly identical. These results suggest that the EM algorithm estimation of haplotype states, frequency, and linkage disequilibrium analysis will be an effective strategy in the CCR2-CCR5 gene region. For genetic epidemiology studies, CCR2-CCR5 allele and haplotype frequencies were determined in African-American (n=30), Hispanic (n=24) and European-American (n=34) populations.

Alleles↗