Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Case-parent triads: estimating single- and double-dose effects of fetal and maternal disease gene haplotypes.

Case-parent triad data are considered a robust basis for studying association between variants of a gene and a disease. Methods evaluating statistical significance of association, like the TDT-test and its extensions, are frequently used. When there are prior hypotheses of a causal effect of the gene under study, however, methods measuring penetrance of alleles or haplotypes as relative risks will be more informative. Log-linear models have been proposed as a flexible tool for such relative risk estimation. We demonstrate an extension of the log-linear model to a natural framework for also estimating effects of multiple alleles or haplotypes, incorporating both single- and double-dose effects. The model also incorporates effects of single- and double-dose maternal haplotypes on a fetus during pregnancy. Unknown phase of haplotypes as well as missing parents are accounted for by the EM algorithm. A number of numerical improvements to maximum likelihood estimation are also implemented to facilitate a larger number of haplotypes. Software for these analyses, HAPLIN, is publicly available through our web site. As an illustration we have re-analyzed data on the MSX1 homeobox-gene on chromosome 4 to show how haplotypes may influence the risk of oral clefts.

Cleft Lip↗

Haplotype-phenotype relationships of paraoxonase-1.

Paraoxonase 1 (PON1) is an enzyme with multiple activities, including detoxification of organophosphates. It is believed to be important in preventing neurotoxic damage and has also been implicated in atherosclerosis. The PON1 gene contains five common polymorphisms, three in the promoter (-909G > C, -162A > G, -108C > T) and two in the coding region (M55L, Q192R) with varying but incomplete linkage disequilibrium. Our previous study showed that functional polymorphisms in PON1 were strongly associated with enzymatic activity in both pregnant women [26-30 weeks of gestation] and neonates. However, there was substantial overlapping of enzyme activities between genotypes. In this study, we investigated whether haplotype (genotype + phase) information would strengthen the genotype-phenotype relationship for PON1. The study consisted of a multiethnic population of 402 mothers and 229 neonates. Haplotypes were imputed by two widely used programs, PHASE and tagSNPs, which yielded very similar results. There were seven haplotypes with a frequency of 5% or higher in at least one ethnic group of the study population. Haplotype composition varied substantially with respect to ethnicity. Haplotypes in Caucasians and African-Americans showed the largest difference, and Caribbean Hispanics seemed to be a mixture of Caucasian and African ancestry. Collectively, the genetic (genotype or haplotype) contribution to PON1 enzymatic activity (measured as phenylacetate hydrolysis) was greater in neonates compared with mothers. Specifically, 16.6% of PON1 variability was explained by genotypes in mothers compared with 30.9% in neonates. Haplotype information offered a slightly increased power in predicting PON1 activity; they explained 35.5% and 19.3% of PON1 variability in neonates and mothers, respectively.

Adult↗

Linkage disequilibrium assessment via log-linear modeling of SNP haplotype frequencies.

Analyses of high-density single-nucleotide polymorphism (SNP) data, such as genetic mapping and linkage disequilibrium (LD) studies, require phase-known haplotypes to allow for the correlation between tightly linked loci. However, current SNP genotyping technology cannot determine phase, which must be inferred statistically. In this paper, we present a new Bayesian Markov chain Monte Carlo (MCMC) algorithm for population haplotype frequency estimation, particularly in the context of LD assessment. The novel feature of the method is the incorporation of a log-linear prior model for population haplotype frequencies. We present simulations to suggest that 1) the log-linear prior model is more appropriate than the standard coalescent process in the presence of recombination (>0.02 cM between adjacent loci), and 2) there is substantial inflation in measures of LD obtained by a "two-stage" approach to the analysis by treating the "best" haplotype configuration as correct, without regard to uncertainty in the recombination process.

Algorithms↗

Polymorphisms in the gene encoding angiotensin I converting enzyme 2 and diabetic nephropathy.

AIMS/HYPOTHESIS: Substantial evidence exists for the involvement of the renin-angiotensin system (RAS) in diabetic nephropathy. Angiotensin I converting enzyme 2 (ACE2), a new component of the RAS, has been implicated in kidney disease, hypertension and cardiac function. Based on this, the aim of the present study was to evaluate whether variations in ACE2 are associated with diabetic nephropathy. MATERIALS AND METHODS: We used a cross-sectional, case-control study design to investigate 823 Finnish type 1 diabetic patients (365 with and 458 without nephropathy). Five single-nucleotide polymorphisms (SNPs) were genotyped using TaqMan technology. Haplotypes were estimated using PHASE software, and haplotype frequency differences were analysed using a chi(2)-test-based tool. RESULTS: None of the ACE2 polymorphisms was associated with diabetic nephropathy, and this finding was supported by the haplotype analysis. The ACE2 polymorphisms were not associated with blood pressure, BMI or HbA(1)c. CONCLUSIONS/INTERPRETATION: In Finnish type 1 diabetic patients, ACE2 polymorphisms are not associated with diabetic nephropathy or any studied risk factor for this complication. Further studies are necessary to assess a minor effect of ACE2.

Adult↗

Haplotype reconstruction from genotype data using Imperfect Phylogeny.

UNLABELLED: Critical to the understanding of the genetic basis for complex diseases is the modeling of human variation. Most of this variation can be characterized by single nucleotide polymorphisms (SNPs) which are mutations at a single nucleotide position. To characterize the genetic variation between different people, we must determine an individual's haplotype or which nucleotide base occurs at each position of these common SNPs for each chromosome. In this paper, we present results for a highly accurate method for haplotype resolution from genotype data. Our method leverages a new insight into the underlying structure of haplotypes that shows that SNPs are organized in highly correlated 'blocks'. In a few recent studies, considerable parts of the human genome were partitioned into blocks, such that the majority of the sequenced genotypes have one of about four common haplotypes in each block. Our method partitions the SNPs into blocks, and for each block, we predict the common haplotypes and each individual's haplotype. We evaluate our method over biological data. Our method predicts the common haplotypes perfectly and has a very low error rate (<2% over the data) when taking into account the predictions for the uncommon haplotypes. Our method is extremely efficient compared with previous methods such as PHASE and HAPLOTYPER. Its efficiency allows us to find the block partition of the haplotypes, to cope with missing data and to work with large datasets. AVAILABILITY: The algorithm is available via a Web server at http://www.calit2.net/compbio/hap/

Algorithms↗

Ancestral genomes, sex, and the population structure of Trypanosoma cruzi.

Acquisition of detailed knowledge of the structure and evolution of Trypanosoma cruzi populations is essential for control of Chagas disease. We profiled 75 strains of the parasite with five nuclear microsatellite loci, 24Salpha RNA genes, and sequence polymorphisms in the mitochondrial cytochrome oxidase subunit II gene. We also used sequences available in GenBank for the mitochondrial genes cytochrome B and NADH dehydrogenase subunit 1. A multidimensional scaling plot (MDS) based in microsatellite data divided the parasites into four clusters corresponding to T. cruzi I (MDS-cluster A), T. cruzi II (MDS-cluster C), a third group of T. cruzi strains (MDS-cluster B), and hybrid strains (MDS-cluster BH). The first two clusters matched respectively mitochondrial clades A and C, while the other two belonged to mitochondrial clade B. The 24Salpha rDNA and microsatellite profiling data were combined into multilocus genotypes that were analyzed by the haplotype reconstruction program PHASE. We identified 141 haplotypes that were clearly distributed into three haplogroups (X, Y, and Z). All strains belonging to T. cruzi I (MDS-cluster A) were Z/Z, the T. cruzi II strains (MDS-cluster C) were Y/Y, and those belonging to MDS-cluster B (unclassified T. cruzi) had X/X haplogroup genotypes. The strains grouped in the MDS-cluster BH were X/Y, confirming their hybrid character. Based on these results we propose the following minimal scenario for T. cruzi evolution. In a distant past there were at a minimum three ancestral lineages that we may call, respectively, T. cruzi I, T. cruzi II, and T. cruzi III. At least two hybridization events involving T. cruzi II and T. cruzi III produced evolutionarily viable progeny. In both events, the mitochondrial recipient (as identified by the mitochondrial clade of the hybrid strains) was T. cruzi II and the mitochondrial donor was T. cruzi III.

Animals↗

Estimation and tests of haplotype-environment interaction when linkage phase is ambiguous.

In the study of complex traits, the utility of linkage analysis and single marker association tests can be limited for researchers attempting to elucidate the complex interplay between a gene and environmental covariates. For these purposes, tests of gene-environment interactions are needed. In addition, recent studies have indicated that haplotypes, which are specific combinations of nucleotides on the same chromosome, may be more suitable as the unit of analysis for statistical tests than single genetic markers. The difficulty with this approach is that, in standard laboratory genotyping, haplotypes are often not directly observable. Instead, unphased marker phenotypes are collected. In this article, we present a method for estimating and testing haplotype-environment interactions when linkage phase is potentially ambiguous. The method builds on the work of Schaid et al. [2002] and is applicable to any trait that can be placed in the generalized linear model framework. Simulations were run to illustrate the salient features of the method. In addition, the method was used to test for haplotype-smoking exposure interaction with data from the Childhood Asthma Management Program.

Algorithms↗

2SNP: scalable phasing based on 2-SNP haplotypes.

2SNP software package implements a new very fast scalable algorithm for haplotype inference based on genotype statistics collected only for pairs of SNPs. This software can be used for comparatively accurate phasing of large number of long genome sequences, e.g. obtained from DNA arrays. As an input 2SNP takes genotype matrix and outputs the corresponding haplotype matrix. On datasets across 79 regions from HapMap 2SNP is several orders of magnitude faster than GERBIL and PHASE while matching them in quality measured by the number of correctly phased genotypes, single-site and switching errors. For example, 2SNP requires 41 s on Pentium 4 2 Ghz processor to phase 30 genotypes with 1381 SNPs (ENm010.7p15:2 data from HapMap) versus GERBIL and PHASE requiring more than a week and admitting no less errors than 2SNP.

Algorithms↗

Family-based tests for associating haplotypes with general phenotype data: application to asthma genetics.

We provide a general purpose family-based testing strategy for associating disease phenotypes with haplotypes when phase may be ambiguous and parental genotype data may be missing. These tests for linkage and association can be used in candidate gene studies with tightly linked markers. Our proposed weighted conditional approach extends the method described in Rabinowitz and Laird to multiple markers. It is attractive because it provides haplotype tests for family-based studies that are efficient and robust to population admixture, phenotype distribution specification, and ascertainment based on phenotypes. It can handle missing parental genotypes and/or missing phase in both offspring and parents. It yields either haplotype-specific (univariate) tests or multi-haplotype (global) tests. This extension has been implemented in the freely available software haplotype FBAT. We used the haplotype FBAT program to test for associations between asthma phenotypes and single nucleotide polymorphisms (SNPs) in the beta-2 adrenergic receptor gene. Whereas no single SNP showed significant association with asthma diagnosis or bronchodilator responsiveness (quantitative trait), a haplotype-based global test found a highly significant association with asthma diagnosis (P value <0.00005) and the measure of bronchodilator responsiveness (P value =0.016).

Algorithms↗

Complex promoter and coding region beta 2-adrenergic receptor haplotypes alter receptor expression and predict in vivo responsiveness.

The human beta(2)-adrenergic receptor gene has multiple single-nucleotide polymorphisms (SNPs), but the relevance of chromosomally phased SNPs (haplotypes) is not known. The phylogeny and the in vitro and in vivo consequences of variations in the 5' upstream and ORF were delineated in a multiethnic reference population and an asthmatic cohort. Thirteen SNPs were found organized into 12 haplotypes out of the theoretically possible 8,192 combinations. Deep divergence in the distribution of some haplotypes was noted in Caucasian, African-American, Asian, and Hispanic-Latino ethnic groups with >20-fold differences among the frequencies of the four major haplotypes. The relevance of the five most common beta(2)-adrenergic receptor haplotype pairs was determined in vivo by assessing the bronchodilator response to beta agonist in asthmatics. Mean responses by haplotype pair varied by >2-fold, and response was significantly related to the haplotype pair (P = 0.007) but not to individual SNPs. Expression vectors representing two of the haplotypes differing at eight of the SNP loci and associated with divergent in vivo responsiveness to agonist were used to transfect HEK293 cells. beta(2)-adrenergic receptor mRNA levels and receptor density in cells transfected with the haplotype associated with the greater physiologic response were approximately 50% greater than those transfected with the lower response haplotype. The results indicate that the unique interactions of multiple SNPs within a haplotype ultimately can affect biologic and therapeutic phenotype and that individual SNPs may have poor predictive power as pharmacogenetic loci.

Base Sequence↗

Maximum-likelihood estimation of molecular haplotype frequencies in a diploid population.

Molecular techniques allow the survey of a large number of linked polymorphic loci in random samples from diploid populations. However, the gametic phase of haplotypes is usually unknown when diploid individuals are heterozygous at more than one locus. To overcome this difficulty, we implement an expectation-maximization (EM) algorithm leading to maximum-likelihood estimates of molecular haplotype frequencies under the assumption of Hardy-Weinberg proportions. The performance of the algorithm is evaluated for simulated data representing both DNA sequences and highly polymorphic loci with different levels of recombination. As expected, the EM algorithm is found to perform best for large samples, regardless of recombination rates among loci. To ensure finding the global maximum likelihood estimate, the EM algorithm should be started from several initial conditions. The present approach appears to be useful for the analysis of nuclear DNA sequences or highly variable loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci.

Algorithms↗

Long-range (17.7 kb) allele-specific polymerase chain reaction method for direct haplotyping of R117H and IVS-8 mutations of the cystic fibrosis transmembrane regulator gene.

Genotyping of genetic polymorphisms is widely used in clinical molecular laboratories to confirm or predict diseases due to single locus mutations. In contrast, very few molecular methods determine the phase or haplotype of two or more mutations that are kilobases apart. In this report, we describe a new method for haplotyping based on long-range allele-specific PCR. Reaction conditions were established to circumvent the incompatibility of using allele-specific primers and a polymerase with proofreading activity. Haplotypes are determined by post-PCR analysis using different detection methods. The clinical application presented here directly determines the phase of two mutations separated by 17.7 kilobases in the cystic fibrosis transmembrane conductance regulator gene. Each mutation, the missense mutation R117H in exon 4 and the 5T polymorphism in intron 8 (IVS-8), have mild phenotypic effect unless they are present on the same chromosome (in cis). If an individual is heterozygous for both R117H and the IVS-8 5T variant, cis/trans testing is required to completely interpret results. The molecular method presented here bypasses the need to perform family studies to establish haplotypes. We propose use of this assay as a reflex clinical test for R117H- 5T-positive samples.

Alleles↗

A note on phasing long genomic regions using local haplotype predictions.

The common approaches for haplotype inference from genotype data are targeted toward phasing short genomic regions. Longer regions are often tackled in a heuristic manner, due to the high computational cost. Here, we describe a novel approach for phasing genotypes over long regions, which is based on combining information from local predictions on short, overlapping regions. The phasing is done in a way, which maximizes a natural maximum likelihood criterion. Among other things, this criterion takes into account the physical length between neighboring single nucleotide polymorphisms. The approach is very efficient and is applied to several large scale datasets and is shown to be successful in two recent benchmarking studies (Zaitlen et al., in press; Marchini et al., in preparation). Our method is publicly available via a webserver at http://research.calit2.net/hap/.

Algorithms↗

Digital genotyping and haplotyping with polymerase colonies.

Polymerase colony (polony) technology amplifies multiple individual DNA molecules within a thin acrylamide gel attached to a microscope slide. Each DNA molecule included in the reaction produces an immobilized colony of double-stranded DNA. We genotype these polonies by performing single base extensions with dye-labeled nucleotides, and we demonstrate the accurate quantitation of two allelic variants. We also show that polony technology can determine the phase, or haplotype, of two single- nucleotide polymorphisms (SNPs) by coamplifying distally located targets on a single chromosomal fragment. We correctly determine the genotype and phase of three different pairs of SNPs. In one case, the distance between the two SNPs is 45 kb, the largest distance achieved to date without separating the chromosomes by cloning or somatic cell fusion. The results indicate that polony genotyping and haplotyping may play an important role in understanding the structure of genetic variation.

DNA↗

A bit about haplotypes: using binary digits to code tightly linked loci.

The advent of recombinant DNA techniques has resulted in the detection of a large number of polymorphic marker loci many of which are not useful for linkage studies because of their low degree of polymorphism. However, when no apparent recombination exists between several closely linked markers, the amount of 'information' available can be significantly increased by establishing haplotypes from those loci; that is, haplotyping increases the number of heterozygotes at marker loci. Haplotyping can be problematic when more than one of the loci involved in the haplotyping are heterozygous and the phase of the haplotypes cannot be inferred from the data. We present a method for recoding the phenotypic marker data for pedigree members that will circumvent this difficulty.

Alleles↗

Characterisation of SNP haplotype structure in chemokine and chemokine receptor genes using CEPH pedigrees and statistical estimation.

Chemokine signals and their cell-surface receptors are important modulators of HIV-1 disease and cancer. To aid future case/control association studies, aim to further characterise the haplotype structure of variation in chemokine and chemokine receptor genes. To perform haplotype analysis in a population-based association study, haplotypes must be determined by estimation, in the absence of family information or laboratory methods to establish phase. Here, test the accuracy of estimates of haplotype frequency and linkage disequilibrium by comparing estimated haplotypes generated with the expectation maximisation (EM) algorithm to haplotypes determined from Centre d'Etude Polymorphisme Humain (CEPH) pedigree data. To do this, they have characterised haplotypes comprising alleles at 11 biallelic loci in four chemokine receptor genes (CCR3, CCR2, CCR5 and CCRL2), which span 150 kb on chromosome 3p21, and haplotyes of nine biallelic loci in six chemokine genes [MCP-1(CCL2), Eotaxin(CCL11), RANTES(CCL5), MPIF-1(CCL23), PARC(CCL18) and MIP-1alpha(CCL3)] on chromosome 17q11-12. Forty multi-generation CEPH families, totalling 489 individuals, were genotyped by the TaqMan 5'-nuclease assay. Phased haplotypes and haplotypes estimated from unphased genotypes were compared in 103 grandparents who were assumed to have mated at random. For the 3p21 single nucleotide polymorphism (SNP) data, haplotypes determined by pedigree analysis and haplotypes generated by the EM algorithm were nearly identical. Linkage disequilibrium, measured by the D' statistic, was nearly maximal across the 150 kb region, with complete disequilibrium maintained at the extremes between CCR3-Y17Y and CCRL2-I243V. D'-values calculated from estimated haplotypes on 3p21 had high concordance with pairwise comparisons between pedigree-phased chromosomes. Conversely, there was less agreement between analyses of haplotype frequencies and linkage disequilibrium using estimated haplotypes when compared with pedigree-phased haplotypes of SNPs on chromosome 17q11-12. These results suggest that, while estimations of haplotype frequency and linkage disequilibrium may be relatively simple in the 3p21 chemokine receptor cluster in population samples, the more complex environment on chromosome 17q11-12 will require a higher resolution haplotype analysis.

Algorithms↗

Haplotype-based linkage disequilibrium mapping via direct data mining.

MOTIVATION: With the availability of large-scale, high-density single-nucleotide polymorphism markers and information on haplotype structures and frequencies, a great challenge is how to take advantage of haplotype information in the association mapping of complex diseases in case-control studies. RESULTS: We present a novel approach for association mapping based on directly mining haplotypes (i.e. phased genotype pairs) produced from case-control data or case-parent data via a density-based clustering algorithm, which can be applied to whole-genome screens as well as candidate-gene studies in small genomic regions. The method directly explores the sharing of haplotype segments in affected individuals that are rarely present in normal individuals. The measure of sharing between two haplotypes is defined by a new similarity metric that combines the length of the shared segments and the number of common alleles around any marker position of the haplotypes, which is robust against recent mutations/genotype errors and recombination events. The effectiveness of the approach is demonstrated by using both simulated datasets and real datasets. The results show that the algorithm is accurate for different population models and for different disease models, even for genes with small effects, and it outperforms some recently developed methods.

Algorithms↗

Haplotype block structure is conserved across mammals.

Genetic variation in genomes is organized in haplotype blocks, and species-specific block structure is defined by differential contribution of population history effects in combination with mutation and recombination events. Haplotype maps characterize the common patterns of linkage disequilibrium in populations and have important applications in the design and interpretation of genetic experiments. Although evolutionary processes are known to drive the selection of individual polymorphisms, their effect on haplotype block structure dynamics has not been shown. Here, we present a high-resolution haplotype map for a 5-megabase genomic region in the rat and compare it with the orthologous human and mouse segments. Although the size and fine structure of haplotype blocks are species dependent, there is a significant interspecies overlap in structure and a tendency for blocks to encompass complete genes. Extending these findings to the complete human genome using haplotype map phase I data reveals that linkage disequilibrium values are significantly higher for equally spaced positions in genic regions, including promoters, as compared to intergenic regions, indicating that a selective mechanism exists to maintain combinations of alleles within potentially interacting coding and regulatory regions. Although this characteristic may complicate the identification of causal polymorphisms underlying phenotypic traits, conservation of haplotype structure may be employed for the identification and characterization of functionally important genomic regions.

Animals↗