Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering↗

Score tests for association between traits and haplotypes when linkage phase is ambiguous.

A key step toward the discovery of a gene related to a trait is the finding of an association between the trait and one or more haplotypes. Haplotype analyses can also provide critical information regarding the function of a gene; however, when unrelated subjects are sampled, haplotypes are often ambiguous because of unknown linkage phase of the measured sites along a chromosome. A popular method of accounting for this ambiguity in case-control studies uses a likelihood that depends on haplotype frequencies, so that the haplotype frequencies can be compared between the cases and controls; however, this traditional method is limited to a binary trait (case vs. control), and it does not provide a method of testing the statistical significance of specific haplotypes. To address these limitations, we developed new methods of testing the statistical association between haplotypes and a wide variety of traits, including binary, ordinal, and quantitative traits. Our methods allow adjustment for nongenetic covariates, which may be critical when analyzing genetically complex traits. Furthermore, our methods provide several different global tests for association, as well as haplotype-specific tests, which give a meaningful advantage in attempts to understand the roles of many different haplotypes. The statistics can be computed rapidly, making it feasible to evaluate the associations between many haplotypes and a trait. To illustrate the use of our new methods, they are applied to a study of the association of haplotypes (composed of genes from the human-leukocyte-antigen complex) with humoral immune response to measles vaccination. Limited simulations are also presented to demonstrate the validity of our methods, as well as to provide guidelines on how our methods could be used.

Algorithms↗

Linkage disequilibrium assessment via log-linear modeling of SNP haplotype frequencies.

Analyses of high-density single-nucleotide polymorphism (SNP) data, such as genetic mapping and linkage disequilibrium (LD) studies, require phase-known haplotypes to allow for the correlation between tightly linked loci. However, current SNP genotyping technology cannot determine phase, which must be inferred statistically. In this paper, we present a new Bayesian Markov chain Monte Carlo (MCMC) algorithm for population haplotype frequency estimation, particularly in the context of LD assessment. The novel feature of the method is the incorporation of a log-linear prior model for population haplotype frequencies. We present simulations to suggest that 1) the log-linear prior model is more appropriate than the standard coalescent process in the presence of recombination (>0.02 cM between adjacent loci), and 2) there is substantial inflation in measures of LD obtained by a "two-stage" approach to the analysis by treating the "best" haplotype configuration as correct, without regard to uncertainty in the recombination process.

Algorithms↗

Estimation and tests of haplotype-environment interaction when linkage phase is ambiguous.

In the study of complex traits, the utility of linkage analysis and single marker association tests can be limited for researchers attempting to elucidate the complex interplay between a gene and environmental covariates. For these purposes, tests of gene-environment interactions are needed. In addition, recent studies have indicated that haplotypes, which are specific combinations of nucleotides on the same chromosome, may be more suitable as the unit of analysis for statistical tests than single genetic markers. The difficulty with this approach is that, in standard laboratory genotyping, haplotypes are often not directly observable. Instead, unphased marker phenotypes are collected. In this article, we present a method for estimating and testing haplotype-environment interactions when linkage phase is potentially ambiguous. The method builds on the work of Schaid et al. [2002] and is applicable to any trait that can be placed in the generalized linear model framework. Simulations were run to illustrate the salient features of the method. In addition, the method was used to test for haplotype-smoking exposure interaction with data from the Childhood Asthma Management Program.

Algorithms↗

Family-based tests for associating haplotypes with general phenotype data: application to asthma genetics.

We provide a general purpose family-based testing strategy for associating disease phenotypes with haplotypes when phase may be ambiguous and parental genotype data may be missing. These tests for linkage and association can be used in candidate gene studies with tightly linked markers. Our proposed weighted conditional approach extends the method described in Rabinowitz and Laird to multiple markers. It is attractive because it provides haplotype tests for family-based studies that are efficient and robust to population admixture, phenotype distribution specification, and ascertainment based on phenotypes. It can handle missing parental genotypes and/or missing phase in both offspring and parents. It yields either haplotype-specific (univariate) tests or multi-haplotype (global) tests. This extension has been implemented in the freely available software haplotype FBAT. We used the haplotype FBAT program to test for associations between asthma phenotypes and single nucleotide polymorphisms (SNPs) in the beta-2 adrenergic receptor gene. Whereas no single SNP showed significant association with asthma diagnosis or bronchodilator responsiveness (quantitative trait), a haplotype-based global test found a highly significant association with asthma diagnosis (P value <0.00005) and the measure of bronchodilator responsiveness (P value =0.016).

Algorithms↗

Complex promoter and coding region beta 2-adrenergic receptor haplotypes alter receptor expression and predict in vivo responsiveness.

The human beta(2)-adrenergic receptor gene has multiple single-nucleotide polymorphisms (SNPs), but the relevance of chromosomally phased SNPs (haplotypes) is not known. The phylogeny and the in vitro and in vivo consequences of variations in the 5' upstream and ORF were delineated in a multiethnic reference population and an asthmatic cohort. Thirteen SNPs were found organized into 12 haplotypes out of the theoretically possible 8,192 combinations. Deep divergence in the distribution of some haplotypes was noted in Caucasian, African-American, Asian, and Hispanic-Latino ethnic groups with >20-fold differences among the frequencies of the four major haplotypes. The relevance of the five most common beta(2)-adrenergic receptor haplotype pairs was determined in vivo by assessing the bronchodilator response to beta agonist in asthmatics. Mean responses by haplotype pair varied by >2-fold, and response was significantly related to the haplotype pair (P = 0.007) but not to individual SNPs. Expression vectors representing two of the haplotypes differing at eight of the SNP loci and associated with divergent in vivo responsiveness to agonist were used to transfect HEK293 cells. beta(2)-adrenergic receptor mRNA levels and receptor density in cells transfected with the haplotype associated with the greater physiologic response were approximately 50% greater than those transfected with the lower response haplotype. The results indicate that the unique interactions of multiple SNPs within a haplotype ultimately can affect biologic and therapeutic phenotype and that individual SNPs may have poor predictive power as pharmacogenetic loci.

Base Sequence↗

Maximum-likelihood estimation of molecular haplotype frequencies in a diploid population.

Molecular techniques allow the survey of a large number of linked polymorphic loci in random samples from diploid populations. However, the gametic phase of haplotypes is usually unknown when diploid individuals are heterozygous at more than one locus. To overcome this difficulty, we implement an expectation-maximization (EM) algorithm leading to maximum-likelihood estimates of molecular haplotype frequencies under the assumption of Hardy-Weinberg proportions. The performance of the algorithm is evaluated for simulated data representing both DNA sequences and highly polymorphic loci with different levels of recombination. As expected, the EM algorithm is found to perform best for large samples, regardless of recombination rates among loci. To ensure finding the global maximum likelihood estimate, the EM algorithm should be started from several initial conditions. The present approach appears to be useful for the analysis of nuclear DNA sequences or highly variable loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci.

Algorithms↗

Digital genotyping and haplotyping with polymerase colonies.

Polymerase colony (polony) technology amplifies multiple individual DNA molecules within a thin acrylamide gel attached to a microscope slide. Each DNA molecule included in the reaction produces an immobilized colony of double-stranded DNA. We genotype these polonies by performing single base extensions with dye-labeled nucleotides, and we demonstrate the accurate quantitation of two allelic variants. We also show that polony technology can determine the phase, or haplotype, of two single- nucleotide polymorphisms (SNPs) by coamplifying distally located targets on a single chromosomal fragment. We correctly determine the genotype and phase of three different pairs of SNPs. In one case, the distance between the two SNPs is 45 kb, the largest distance achieved to date without separating the chromosomes by cloning or somatic cell fusion. The results indicate that polony genotyping and haplotyping may play an important role in understanding the structure of genetic variation.

DNA↗

A bit about haplotypes: using binary digits to code tightly linked loci.

The advent of recombinant DNA techniques has resulted in the detection of a large number of polymorphic marker loci many of which are not useful for linkage studies because of their low degree of polymorphism. However, when no apparent recombination exists between several closely linked markers, the amount of 'information' available can be significantly increased by establishing haplotypes from those loci; that is, haplotyping increases the number of heterozygotes at marker loci. Haplotyping can be problematic when more than one of the loci involved in the haplotyping are heterozygous and the phase of the haplotypes cannot be inferred from the data. We present a method for recoding the phenotypic marker data for pedigree members that will circumvent this difficulty.

Alleles↗

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9&#xa0;Gb and represented 23 pairs of chromosomes with an NG50 of 150&#xa0;Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans↗

The accuracy of statistical methods for estimation of haplotype frequencies: an example from the CD4 locus.

Haplotype analysis has become increasingly important for the study of human disease as well as for reconstruction of human population histories. Computer programs have been developed to estimate haplotype frequencies statistically from marker phenotypes in unrelated individuals. However, there currently are few empirical reports on the accuracy of statistical estimates that must infer linkage phase. We have analyzed haplotypes at the CD4 locus on chromosome 12 that consist of a short tandem-repeat polymorphism and an Alu insertion/deletion polymorphism located 9.8 kb apart, in 398 individuals from 10 geographically diverse sub-Saharan African populations. Haplotype frequency estimates obtained using gene counting based on molecularly haplotyped (phase-known) data were compared with haplotype frequency estimates obtained using the expectation-maximization algorithm. We show that the estimated frequencies of common haplotypes do not differ significantly with the use of phase-known versus phase-unknown data. However, rare haplotypes are occasionally miscalled when their presence/absence must be inferred. Thus, for those research questions for which the common haplotypes are most important, frequency estimates based on the phase-unknown marker-typing results from unrelated individuals will be sufficient. However, in cases where knowledge of rare haplotypes is critical, molecular haplotyping will be necessary to determine linkage phase unambiguously.

Africa South of the Sahara↗

Biallelic VPS41 Variants in Autosomal Recessive Spinocerebellar Ataxia 29 Resolved by Long-Read Sequencing and RNA Analysis.

BACKGROUND: Biallelic variants in VPS41, encoding a subunit of the HOPS complex, cause autosomal recessive spinocerebellar ataxia 29 (SCAR29), a rare neurodevelopmental disorder with an incompletely defined phenotypic and molecular spectrum. METHODS: We investigated a 24-year-old man with cerebellar ataxia, hypotonia, and intellectual disability. Exome sequencing identified four candidate VPS41 variants. Because maternal DNA was unavailable, long-read genome sequencing was performed to determine allelic configuration, followed by RNA and protein analyses. RESULTS: In addition to typical SCAR29 features, the patient showed previously unreported findings, including swan-neck deformities and pes cavus. Long-read genome sequencing demonstrated that two VPS41 variants were in trans. RNA analysis revealed distinct splicing consequences: one allele produced an out-of-frame transcript predicted to undergo nonsense-mediated decay, whereas the other generated an in-frame exon-skipped transcript. These complementary defects reduced VPS41 expression at both transcript and protein levels, supporting pathogenicity and variant reclassification. CONCLUSION: Our findings expand the phenotypic spectrum of VPS41-related disease and highlight the value of long-read allelic resolution in clarifying pathogenic mechanisms in rare genetic disorders.

Humans↗

Haplotype-resolved reconstruction and functional interrogation of cancer karyotypes.

Complex karyotype changes are widespread in cancer genomes. A major gap in cancer genome characterization is the resolution of rearranged chromosomes with chromosome-length continuity. Here, we describe a two-tiered approach to determine the segmental composition of rearranged chromosomes with haplotype resolution. First, we present refLinker, a bioinformatic method for robust determination of chromosomal haplotypes using cancer Hi-C data. By contrast with existing methods, refLinker is insensitive to the presence of large-scale DNA deletions, duplications, and high-level amplification in cancer genomes. Second, we demonstrate a computational strategy to determine the segmental structure of rearranged chromosomes using haplotype-specific Hi-C contacts. We apply these methods to breast cancer genomes and provide direct evidence for long-range transcriptional changes associated with rearrangements of the inactive X chromosome. Together, these results highlight refLinker's broad utility for studying the functional consequences of chromosomal rearrangements.

Humans↗

A novel method for across-chromosome phasing without relative data.

MOTIVATION: Across-chromosome phasing identifies which haplotypes of different chromosomes come from the same parent. This differs from within-chromosome phasing, which uses linkage disequilibrium patterns to determine which alleles were co-inherited within each chromosome but does not match haplotypes across different chromosomes. While across-chromosome phasing can be conducted using genotypes from parents or close relatives, current methods perform poorly for samples of unrelated individuals. Here, we introduce a novel approach for across-chromosome phasing that employs a window-based SNP-similarity metric, eliminating the need for data from close relatives or detection of identical-by-descent haplotypes. RESULTS: Using UK Biobank offspring with both parents genotyped as a gold standard, we evaluated the performance of our method by phasing the offspring without using parental data. In genomic data with no within-chromosome phase errors, our algorithm achieved a mean across-chromosome phasing accuracy of 95%, with 53% of individuals phased perfectly. When data was pre-phased computationally using a standard within-chromosome phasing algorithm, mean accuracy for across-chromosome phasing dropped to 83.1%. Thus, our method is limited primarily by the accuracy of within-chromosome phasing accuracy and can approach near-perfect across-chromosome phasing accuracy as within-chromosome phasing accuracy improves. AVAILABILITY AND IMPLEMENTATION: The implementation was executed within a multi-node computational environment of University of Colorado Boulder Research Computing (Blanca Cluster: https://www.colorado.edu/rc/resources/blanca), employing parallelization techniques in the C programming language. The source code has been made publicly accessible online at https://github.com/emmanuelsapin/AcrossChromosomesPhasing, thereby facilitating reproducibility of the results for researchers with authorized access to the UK Biobank dataset.

Algorithms↗

Haplotypic analyses of the IGF2-INS-TH gene cluster in relation to cardiovascular risk traits.

The IGF2-INS-TH genomic region has been implicated in various common disorders including the metabolic syndrome, type 2 diabetes and coronary heart disease (CHD). Here we present detailed haplotype analysis of 2743 males 51-62 years old in relation to body weight and composition, blood pressure (BP) and plasma triglycerides (TG). Use of the total data set was complicated by the number of loci typed, missing data, multi-allelic markers and continuous trait phenotypes. Different algorithms and subsets of the data were analysed using the programmes haplotype trend regression, haplo.score, evolutionary-based haplotype analysis package and Phase, in conjunction with SPSS. Ten haplotypes designated in frequency order *1(20.0%) to *10(3.4%) represented 89% of all haplotypes. Haplotype *5 protected against obesity. Haplotype *4 carriers exhibited elevated BP and fat mass, haplotype *6 was associated with raised plasma TG levels. Haplotype *8 also showed similar magnitude effects as *4. These cohort trait analyses and detailed haplotypic analyses enable integration with published case data. Haplotypes *4, *6 and *8 are the only INS VNTR class III-bearing haplotypes, although differing in flanking haplotype, whereas *5 displays unique features in all three genes (with significant commonality with type 1 diabetes-predisposition haplotypes). We propose that long repeat insertion in the insulin gene promoter ('class III'), reported to result in low insulin production, predisposes to the metabolic syndrome features of elevated BP, fat mass or TG level, therefore appearing more frequently in type 2 diabetic, polycystic ovary syndrome and CHD cases. The functional element(s) of *5 for weight-lowering could reside in any of the three genes.

Aged↗

A unified stepwise regression procedure for evaluating the relative effects of polymorphisms within a gene using case/control or family data: application to HLA in type 1 diabetes.

A stepwise logistic-regression procedure is proposed for evaluation of the relative importance of variants at different sites within a small genetic region. By fitting statistical models with main effects, rather than modeling the full haplotype effects, we generate tests, with few degrees of freedom, that are likely to be powerful for detecting primary etiological determinants. The approach is applicable to either case/control or nuclear-family data, with case/control data modeled via unconditional and family data via conditional logistic regression. Four different conditioning strategies are proposed for evaluation of effects at multiple, closely linked loci when family data are used. The first strategy results in a likelihood that is equivalent to analysis of a matched case/control study with each affected offspring matched to three pseudocontrols, whereas the second strategy is equivalent to matching each affected offspring with between one and three pseudocontrols. Both of these strategies require you be able to infer parental phase (i.e., those haplotypes present in the parents). Families in which phase cannot be determined must be discarded, which can considerably reduce the effective size of a data set, particularly when large numbers of loci that are not very polymorphic are being considered. Therefore, a third strategy is proposed in which knowledge of parental phase is not required, which allows those families with ambiguous phase to be included in the analysis. The fourth and final strategy is to use conditioning method 2 when parental phase can be inferred and to use conditioning method 3 otherwise. The methods are illustrated using nuclear-family data to evaluate the contribution of loci in the HLA region to the development of type 1 diabetes.

Alleles↗

An MLH1 haplotype is over-represented on chromosomes carrying an HNPCC predisposing mutation in MLH1.

BACKGROUND: The mismatch repair gene, MLH1, appears to occur as two main haplotypes at least in white populations. These are referred to as A and G types with reference to the A/G polymorphism at IVS14-19. On the basis of preliminary experimental data, we hypothesised that deviations from the expected frequency of these two haplotypes could exist in carriers of disease associated MLH1 germline mutations. METHODS: We assembled a series (n=119) of germline MLH1 mutation carriers in whom phase between the haplotype and the mutation had been conclusively established. Controls, without cancer, were obtained from each contributing centre. Cases and controls were genotyped for the polymorphism in IVS14. RESULTS: Overall, 66 of 119 MLH1 mutations occurred on a G haplotype (55.5%), compared with 315 G haplotypes on 804 control chromosomes (39.2%, p=0.001). The odds ratio (OR) of a mutation occurring on a G rather than an A haplotype was 1.93 (95% CI 1.29 to 2.91). When we compared the haplotype frequencies in mutation bearing chromosomes carried by people of different nationalities with those seen in pooled controls, all groups showed a ratio of A/G haplotypes that was skewed towards G, except the Dutch group. On further analysis of the type of each mutation, it was notable that, compared with control frequencies, deletion and substitution mutations were preferentially represented on the G haplotype (p=0.003 and 0.005, respectively). CONCLUSION: We have found that disease associated mutations in MLH1 appear to occur more often on one of only two known ancient haplotypes. The underlying reason for this observation is obscure, but it is tempting to suggest a possible role of either distant regulatory sequences or of chromatin structure influencing access to DNA sequence. Alternatively, differential behaviour of otherwise similar haplotypes should be considered as prime areas for further study.

Adaptor Proteins, Signal Transducing↗

Analysis and exploration of the use of rule-based algorithms and consensus methods for the inferral of haplotypes.

The difficulty of experimental determination of haplotypes from phase-unknown genotypes has stimulated the development of nonexperimental inferral methods. One well-known approach for a group of unrelated individuals involves using the trivially deducible haplotypes (those found in individuals with zero or one heterozygous sites) and a set of rules to infer the haplotypes underlying ambiguous genotypes (those with two or more heterozygous sites). Neither the manner in which this "rule-based" approach should be implemented nor the accuracy of this approach has been adequately assessed. We implemented eight variations of this approach that differed in how a reference list of haplotypes was derived and in the rules for the analysis of ambiguous genotypes. We assessed the accuracy of these variations by comparing predicted and experimentally determined haplotypes involving nine polymorphic sites in the human apolipoprotein E (APOE) locus. The eight variations resulted in substantial differences in the average number of correctly inferred haplotype pairs. More than one set of inferred haplotype pairs was found for each of the variations we analyzed, implying that the rule-based approach is not sufficient by itself for haplotype inferral, despite its appealing simplicity. Accordingly, we explored consensus methods in which multiple inferrals for a given ambiguous genotype are combined to generate a single inferral; we show that the set of these "consensus" inferrals for all ambiguous genotypes is more accurate than the typical single set of inferrals chosen at random. We also use a consensus prediction to divide ambiguous genotypes into those whose algorithmic inferral is certain or almost certain and those whose less certain inferral makes molecular inferral preferable.

Algorithms↗