Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A generalization of the transmission/disequilibrium test for uncertain-haplotype transmission.

A new transmission/disequilibrium-test statistic is proposed for situations in which transmission is uncertain. Such situations arise when transmission of a multilocus marker haplotype is considered, since haplotype phase is often unknown in a substantial number of instances. Even for single-locus markers, transmission is uncertain if one or both parents are missing. In both these situations, uncertainty may be reduced by the typing of further siblings, whose disease status may be unaffected or unknown. The proposed test is a score test based on a partial score function that omits the terms most influenced by hidden population stratification.

Adult↗

Haplotypic analyses of the aldosterone synthase gene CYP11B2 associated with stage-2 hypertension in northern Han Chinese.

The objective of this study was to investigate the association of polymorphisms in the aldosterone synthase gene CYP11B2 (-344T/C, Lys173/Arg, and an intronic conversion [IC]) with stage-2 hypertension in northern Han Chinese. A total of 503 hypertensives and their age-, gender-, and area-matched controls were included in this study. The female hypertensives had significantly higher frequencies of the -344T, 173Lys, and IC-conversion alleles (p = 0.002, 0.002, and 0.014, respectively). The estimated frequency of haplotype composed of the -344T, 173Lys, and IC-conversion alleles (haplotype 4) was significantly higher in the female hypertensives compared with their controls (p = 0.007). Using a multivariate score test, we found that haplotype 4 remained associated with female hypertension after the adjustment for covariates (p = 0.003), while the haplotype 3 of T-Arg-WT showed a protective effect both in the males and in the females (p = 0.03 and 0.006, respectively). The odds ratio for haplotype phase of 4-4 was 2.60 (95% CI, 1.21-5.58) and for 3-3, 0.20 (95% CI, 0.03-0.77). These results indicate that the Lys173 and the IC-conversion allele of the CYP11B2 gene confer an increased risk for stage-2 hypertension in northern Han Chinese women.

Asian People↗

Rhesus blood group haplotype determination by nanopore sequencing and adaptive sampling enables the precise determination of complex allele combinations that could not be accurately determined by standard methods.

BACKGROUND: Patients with chronic transfusion needs such as those with sickle cell disease face a high risk of developing antibodies against high-prevalence antigens in the RH blood group system, complicating transfusion therapy and potentially necessitating stem cell transplantation. Molecular characterization of the RH system is hindered by hybrid alleles and high sequence homology between RHD and RHCE, limiting the effectiveness of conventional short-read sequencing. STUDY DESIGN AND METHODS: We analyzed 11 control and 20 patient samples, some of which could not be reliably genotyped by standard methods. RESULTS: Nanopore sequencing with adaptive sampling enables targeted, amplification-free long-read sequencing of the RH locus, resolving homologous and complex hybrid structures and enabling complete haplotype phasing for all samples, including samples that could not be accurately determined by standard methods like serology and short-read sequencing. Four new alleles were identified and for 13 out of 20 patients the results led to a change in the transfusion regimen. DISCUSSION: These findings show that nanopore sequencing with adaptive sampling allows unambiguous genotyping of the RH system, improves detection of complex variants, and supports better-matched transfusion strategies for chronically transfused patients.

Rh-Hr Blood-Group System↗

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA↗

Bayesian modelling of multivariate quantitative traits using seemingly unrelated regressions.

We investigate a Bayesian approach to modelling the statistical association between markers at multiple loci and multivariate quantitative traits. In particular, we describe the use of Bayesian Seemingly Unrelated Regressions (SUR) whereby genotypes at the different loci are allowed to have non-simultaneous effects on the phenotypes considered with residuals from each regression assumed correlated. We present results from simulations showing that, under rather general conditions that are likely to hold in real situations, the Bayesian SUR approach has increased probability of selecting the true model compared to univariate analyses. Finally, we apply our methods to data from subjects genotyped for 12 SNPs in the apolipoprotein E (APOE) gene. Phenotypes relate to response to treatment with atorvastatin and include changes in total cholesterol, low-density lipoprotein cholesterol, and triglycerides. Missing genotype data are naturally accommodated in our Bayesian framework by imputing them using a nested haplotype phasing algorithm.

Algorithms↗

Individualized antisense oligonucleotides for SCN2A-related developmental epileptic encephalopathy.

SCN2A variants are among the most common genetic causes of developmental and epileptic encephalopathies (DEEs), which can present with uncontrolled seizures at birth and account for 1-2% of all epileptic encephalopathies. A substantial fraction of causal variants are gain-of-function or mixed-function variants associated with increased channel open probability or greater sodium current flux. Here two parallel n = 1 clinical studies were conducted in two patients (9-year-old and 14-year-old boys) with SCN2A-related DEE. Individualized allele-selective antisense oligonucleotides (ASOs) were designed to target heterozygous intronic single-nucleotide polymorphisms (SNPs) for decreased expression of mutant SCN2A transcript while preserving the wild-type copy. Primary endpoints included quantitative change from baseline in seizure frequency and neurodevelopment, including motor scores. Efficacy measures were also individualized to each patient's phenotype, including refractory seizures, developmental delay, autism spectrum disorder, choreoathetosis and gastrointestinal dysfunction. Patients experienced a reduction in seizure frequency (26% and 90% in the two patients, respectively), decreased use of concomitant medications and improvement in neurodevelopmental skills. Both ASOs were well tolerated, with no ASO-related serious adverse events. Continued long-term follow-up of these preliminary positive safety and efficacy findings is needed to confirm the disease-modifying potential of these ASOs. Haplotype phasing in a separate cohort of infants with SCN2A-related disorder (SCN2A-RD), diagnosed by rapid whole-genome sequencing, identified 16% of patients with compatible SNPs. These data provide a pathway from n = 1 to n of more patients with SCN2A-RD and other monogenic disorders. ClinicalTrials.gov registration: NCT06314490 .

Adolescent↗

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Tubers are wide-spread across angiosperms with some species such as Solanum tuberosum L. (potato) serving as a vital crop for human consumption. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of Flowering Time, that functions as tuberigen, the equivalent of florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of Atlantic, a tetraploid cultivar. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129,218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Journal Article↗

High-resolution characterization of linkage disequilibrium structure and selection of tagging single nucleotide polymorphisms: application to the cholesteryl ester transfer protein gene.

Full characterization of intragenic variation may improve candidate gene associations. This study selected tagging (t) single nucleotide polymorphisms (SNPs) to comprehensively represent genetic variability in the cholesteryl ester transfer protein (CETP) gene. Nineteen SNPs were identified in 50 unrelated individuals in the SNP discovery phase, and 13 intronic SNPs were added from the literature. These 32 SNPs were genotyped in 339 apparently healthy individuals and 190 coronary artery disease (CAD) patients. Using phased haplotypes, linkage disequilibrium (LD) structure was characterized and tSNPs selected using a principal component analysis (PCA) method. In healthy individuals, seven LD groups were identified that accounted for 93.4% of the observed genetic variation. These LD groups highlighted a complex LD structure for CETP, including both recombination and mutation, and eleven tSNPs were selected. Among CAD patients the results were essentially the same. Results from PCA using diploid genotype data were reasonably comparable. Finally, the selected tSNPs successfully represented the association evidence discovered for all of the other SNPs studied. This study provides an optimal set of tSNPs for association analyses of CETP. The observed complexity of LD structure highlights the importance of using methods, such as PCA, that allow for multiple dynamics in intragenic LD structure.

Aged↗

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Potato (Solanum tuberosum L.) is cultivated for its tubers, which serve as a major crop. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of flowering time, that functions as a tuberigen, the equivalent of a florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of the tetraploid potato cultivar, Atlantic. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129 218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Solanum tuberosum↗

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering↗

Score tests for association between traits and haplotypes when linkage phase is ambiguous.

A key step toward the discovery of a gene related to a trait is the finding of an association between the trait and one or more haplotypes. Haplotype analyses can also provide critical information regarding the function of a gene; however, when unrelated subjects are sampled, haplotypes are often ambiguous because of unknown linkage phase of the measured sites along a chromosome. A popular method of accounting for this ambiguity in case-control studies uses a likelihood that depends on haplotype frequencies, so that the haplotype frequencies can be compared between the cases and controls; however, this traditional method is limited to a binary trait (case vs. control), and it does not provide a method of testing the statistical significance of specific haplotypes. To address these limitations, we developed new methods of testing the statistical association between haplotypes and a wide variety of traits, including binary, ordinal, and quantitative traits. Our methods allow adjustment for nongenetic covariates, which may be critical when analyzing genetically complex traits. Furthermore, our methods provide several different global tests for association, as well as haplotype-specific tests, which give a meaningful advantage in attempts to understand the roles of many different haplotypes. The statistics can be computed rapidly, making it feasible to evaluate the associations between many haplotypes and a trait. To illustrate the use of our new methods, they are applied to a study of the association of haplotypes (composed of genes from the human-leukocyte-antigen complex) with humoral immune response to measles vaccination. Limited simulations are also presented to demonstrate the validity of our methods, as well as to provide guidelines on how our methods could be used.

Algorithms↗

How to quantify information loss due to phase ambiguity in haplotype case-control studies.

Assigning haplotypes in a case-control study is a challenging problem. We proposed a method to quantify the information loss due to missing phase information. We determined which individuals were responsible for the information loss, and calculated how much information could be gained when the ambiguous individuals could be resolved by adding additional parental information.

Case-Control Studies↗

Inferring haplotypes at the NAT2 locus: the computational approach.

BACKGROUND: Numerous studies have attempted to relate genetic polymorphisms within the N-acetyltransferase 2 gene (NAT2) to interindividual differences in response to drugs or in disease susceptibility. However, genotyping of individuals single-nucleotide polymorphisms (SNPs) alone may not always provide enough information to reach these goals. It is important to link SNPs in terms of haplotypes which carry more information about the genotype-phenotype relationship. Special analytical techniques have been designed to unequivocally determine the allocation of mutations to either DNA strand. However, molecular haplotyping methods are labour-intensive and expensive and do not appear to be good candidates for routine clinical applications. A cheap and relatively straightforward alternative is the use of computational algorithms. The objective of this study was to assess the performance of the computational approach in NAT2 haplotype reconstruction from phase-unknown genotype data, for population samples of various ethnic origin. RESULTS: We empirically evaluated the effectiveness of four haplotyping algorithms in predicting haplotype phases at NAT2, by comparing the results with those directly obtained through molecular haplotyping. All computational methods provided remarkably accurate and reliable estimates for NAT2 haplotype frequencies and individual haplotype phases. The Bayesian algorithm implemented in the PHASE program performed the best. CONCLUSION: This investigation provides a solid basis for the confident and rational use of computational methods which appear to be a good alternative to infer haplotype phases in the particular case of the NAT2 gene, where there is near complete linkage disequilibrium between polymorphic markers.

Algorithms↗

Tests of trait-haplotype association when linkage phase is ambiguous, appropriate for matched case-control and cohort studies with competing risks.

The impact of competing risks on tests of association between disease and haplotypes has been largely ignored. We consider situations in which linkage phase is ambiguous and show that tests for disease-haplotype association can lead to rejection of the null hypothesis, even when true, with more than the nominal 5 per cent frequency. This problem tends to occur if a haplotype is associated with overall mortality, even if the haplotype is not associated with disease risk. A small simulation study illustrates the magnitude of bias (high type I error rate) in the context of a cohort study in which a modest number of disease cases (about 350) occur over time. The bias remains even if the score test is based on a logistic model that includes age as a covariate. For cohort studies, we propose a new test based on a modification of the proportional hazards model and for case-control studies, a test based on a conditional likelihood that have the correct size under the null even in the presence of competing risks, and that can be used when haplotype is ambiguous.

Aged↗

Identification and functional analysis of common human flavin-containing monooxygenase 3 genetic variants.

Flavin-containing monooxygenases (FMOs) are important for the disposition of many therapeutics, environmental toxicants, and nutrients. FMO3, the major adult hepatic FMO enzyme, exhibits significant interindividual variation. Eighteen FMO3 single-nucleotide polymorphism (SNP) frequencies were determined in 202 Hispanics (Mexican descent), 201 African Americans, and 200 non-Latino whites. Using expressed recombinant enzyme with methimazole, trimethylamine, sulindac, and ethylenethiourea, the novel structural variants FMO3 E24D and K416N were shown to cause modest changes in catalytic efficiency, whereas a third novel variant, FMO3 N61K, was essentially devoid of activity. The latter variant was present at an allelic frequency of 5.2% in non-Latino whites and 3.5% in African Americans, but it was absent in Hispanics. Inferring haplotypes using PHASE, version 2.1, the greatest haplotype diversity was observed in African Americans followed by non-Latino whites and Hispanics. Haplotype 2A and 2B, consisting of a hypermorphic promoter SNP cluster (-2650C>G, -2543T>A, and -2177G>C) in linkage with synonymous structural variants was inferred at a frequency of 27% in the Hispanic population, but only 5% in non-Latino whites and African Americans. This same promoter SNP cluster in linkage with one or more hypomorphic structural variant also was inferred in multiple haplotypes at a total frequency of 5.6% in the African-American study group but less than 1% in the other two groups. The sum frequencies of the hypomorphic haplotypes H3 [15,167G>A (E158K)], H5B [-2650C>G, 15,167G>A (E158K), 21,375C>T (N285N), 21,443A>G (E308G)], and H6 [15,167G>A (E158K), 21,375C>T (N285N)] was 28% in Hispanics, 23% in non-Latino whites, and 24% in African Americans.

Black or African American↗

Comparison of the accuracy of methods of computational haplotype inference using a large empirical dataset.

BACKGROUND: Analyses of genetic data at the level of haplotypes provide increased accuracy and power to infer genotype-phenotype correlations and evolutionary history of a locus. However, empirical determination of haplotypes is expensive and laborious. Therefore, several methods of inferring haplotypes from unphased genotypic data have been proposed, but it is unclear how accurate each of the methods is or which methods are superior. The accuracy of some of the leading methods of computational haplotype inference (PL-EM, Phase, SNPHAP, Haplotyper) are compared using a large set of 308 empirically determined haplotypes based on 15 SNPs, among which 36 haplotypes were observed to occur. This study presents several advantages over many previous comparisons of haplotype inference methods: a large number of subjects are included, the number of known haplotypes is much smaller than the number of chromosomes surveyed, a range in values of linkage disequilibrium, presence of rare SNP alleles, and considerable dispersion in the frequencies of haplotypes. RESULTS: In contrast to some previous comparisons of haplotype inference methods, there was very little difference in the accuracy of the various methods in terms of either assignment of haplotypes to individuals or estimation of haplotype frequencies. Although none of the methods inferred all of the known haplotypes, the assignment of haplotypes to subjects was about 90% correct for individuals heterozygous for up to three SNPs and was about 80% correct for up to five heterozygous sites. All of the methods identified every haplotype with a frequency above 1%, and none assigned a frequency above 1% to an incorrect haplotype. CONCLUSIONS: All of the methods of haplotype inference have high accuracy and one can have confidence in inferences made by any one of the methods. The ability to identify even rare (>/= 1%) haplotypes is reassuring for efforts to identify haplotypes that contribute to disease in a significant proportion of a population. Assignment of haplotypes is relatively accurate among subjects heterozygous for up to 5 sites, and this might be the largest number of SNPs for which one should define haplotype blocks or have confidence in haplotype assignments.

Algorithms↗

Linkage disequilibrium mapping via cladistic analysis of phase-unknown genotypes and inferred haplotypes in the Genetic Analysis Workshop 14 simulated data.

We recently described a method for linkage disequilibrium (LD) mapping, using cladistic analysis of phased single-nucleotide polymorphism (SNP) haplotypes in a logistic regression framework. However, haplotypes are often not available and cannot be deduced with certainty from the unphased genotypes. One possible two-stage approach is to infer the phase of multilocus genotype data and analyze the resulting haplotypes as if known. Here, haplotypes are inferred using the expectation-maximization (EM) algorithm and the best-guess phase assignment for each individual analyzed. However, inferring haplotypes from phase-unknown data is prone to error and this should be taken into account in the subsequent analysis. An alternative approach is to analyze the phase-unknown multilocus genotypes themselves. Here we present a generalization of the method for phase-known haplotype data to the case of unphased SNP genotypes. Our approach is designed for high-density SNP data, so we opted to analyze the simulated dataset. The marker spacing in the initial screen was too large for our method to be effective, so we used the answers provided to request further data in regions around the disease loci and in null regions. Power to detect the disease loci, accuracy in localizing the true site of the locus, and false-positive error rates are reported for the inferred-haplotype and unphased genotype methods. For this data, analyzing inferred haplotypes outperforms analysis of genotypes. As expected, our results suggest that when there is little or no LD between a disease locus and the flanking region, there will be no chance of detecting it unless the disease variant itself is genotyped.

Chromosome Mapping↗

A simple, bead-based approach for multi-SNP molecular haplotyping.

Single nucleotide polymorphisms (SNPs) within a gene region have often been studied to determine their effect on phenotype. Although a single base pair change can produce a phenotypic change, phenotype is often influenced by the presence of multiple polymorphisms and their relative positions within a given region. For example, if multiple changes occur in a promoter region, how they influence gene expression will depend on their cis/trans configuration. As such, it is essential to consider the haplotype, or the alignment of multiple SNP alleles on each chromosome when attempting to associate genomic changes with phenotype. Unfortunately, no method of high-throughput molecular haplotyping of multiple SNPs currently exists. In response to this unmet need, we have developed an inexpensive, reliable bead-based capture-based haplotyping (CBH) assay to determine the phase, or haplotype, of multiple SNP alleles in a high-throughput manner. The CBH assay requires minimal setup and handling, requires no centrifugation steps and can be performed in <1 h. Data collection is performed via flow cytometry and the assay yields plus/minus results allowing for automated calling by a simple computer application. We will present data demonstrating the molecular haplotyping of 11 SNPs within exon 2 of the N-acetyltransferase-2 (NAT2) gene, which expresses an important drug-metabolizing enzyme. This assay has applications in diagnostic testing, promoter analysis, association studies and pharmacogenetic analysis.

Arylamine N-Acetyltransferase↗