Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Haplotype interaction analysis of unlinked regions.

Genetically complex diseases are caused by interacting environmental factors and genes. As a consequence, statistical methods that consider multiple unlinked genomic regions simultaneously are desirable. Such consideration, however, may lead to a vast number of different high-dimensional tests whose appropriate analysis pose a problem. Here, we present a method to analyze case-control studies with multiple SNP data without phase information that considers gene-gene interaction effects while correcting appropriately for multiple testing. In particular, we allow for interactions of haplotypes that belong to different unlinked regions, as haplotype analysis often proves to be more powerful than single marker analysis. In addition, we consider different marker combinations at each unlinked region. The multiple testing issue is settled via the minP approach; the P value of the "best" marker/region configuration is corrected via Monte-Carlo simulations. Thus, we do not explicitly test for a specific pre-defined interaction model, but test for the global hypothesis that none of the considered haplotype interactions shows association with the disease. We carry out a simulation study for case-control data that confirms the validity of our approach. When simulating two-locus disease models, our test proves to be more powerful than association methods that analyze each linked region separately. In addition, when one of the tested regions is not involved in the etiology of the disease, only a small amount of power is lost with interaction analysis as compared to analysis without interaction. We successfully applied our method to a real case-control data set with markers from two genes controlling a common pathway. While classical analysis failed to reach significance, we obtained a significant result even after correction for multiple testing with our proposed haplotype interaction analysis. The method described here has been implemented in FAMHAP.

Algorithms↗

A new strategy for mannose-binding lectin gene haplotyping.

The mannose-binding lectin 2 (MBL2) gene is polymorphic and codes for a protein with an important role in the innate immune response, whose variants have been associated with a great number of diseases. Point variations have been described in the 5' regulatory region at positions -550 (MBL2*H or *L) and -221 (*X or *Y), in the 5' untranslated sequence at position +4 (*P or *Q), and in the coding sequence of exon 1 at codons 52, 54, and 57 (MBL2*A or D, A or B, and A or C, respectively). These can be in cis or in trans configuration. The different haplotypes influence the immunological phenotype of the individual, which makes MBL2 haplotyping very important. Previously described MBL2-typing methods do not present adequate haplotype resolution or are too complex and costly. We have developed a new MBL2-typing strategy that is economical and renders rapid and reliable results without ambiguities. We typed 202 individuals of European, 32 of African, and 16 of Oriental descent. Only five to six reactions from 10 possible PCR-SSPs (sequence-specific polymerase chain reactions) were sufficient to genotype one individual unambiguously. The reactions were specific for amplification of the variants located upstream of the coding sequence. The results were associated to the results of hybridizations of the amplified products with eight sequence-specific oligonucleotide probes (SSOP). The strategy led to identification of eight alleles: MBL2*HYPA, HYPD, LYPA, LYPB, LYPD, LYQA, LYQC, and LXPA. Their frequencies in each of the groups were similar to those of other populations studied to date, with MBL2*LYPD (g.[-550G>C; -221C>G; 4T>C; 223C>T; 230A>G; 239A>G]) being novel. All samples were found to be in Hardy-Weinberg equilibrium.

Acute-Phase Proteins↗

[How about the uncertainty in the haplotypes in the population-based KORA studies?].

In the KORA surveys, numerous candidate genes in the context of type 2 diabetes, myocardial infarction, atherosclerosis or obesity are under investigation. Current focus is on genotyping single nucleotide polymorphism (SNPs). Haplotypes are also of increasing interest: haplotypes are combinations of alleles within a certain section of one chromosome. Analysing haplotypes in genetic association studies is often more efficient than studying the SNPs separately. A statistical problem in this context is the reconstruction of the phase: genotyping the SNPs determines the alleles of an individual at one particular locus of the DNA, but does not reveal which allele is located on which one of the two chromosomes. This information is required when talking about haplotypes. There are statistical approaches to identify the most likely two haplotypes of an individual given the genotypes. However, a certain error in prognosis is unavoidable. There are also errors in the genotypes. These errors are assumed to be small for one SNP but can accumulate over the SNPs involved in one haplotype and thus can induce further uncertainty in the haplotype. It is therefore the aim of our project to quantify the uncertainties in the haplotypes particularly for genes investigated in the KORA surveys. We conduct computer simulations based on the haplotypes and their frequencies observed in the KORA individuals and compare the results with simulations based on mathematical modelling of the evolutionary process ("coalescent models"). The uncertainties in the haplotypes have an impact on the search for association between genes and disease: an association may not be detected as the haplotype uncertainty obscures the haplotype frequency differences between cases and controls. It is a further aim of our project to elucidate the extent of this problem and to develop strategies for reducing it.

Adult↗

Design and sample-size considerations in the detection of linkage disequilibrium with a disease locus.

The presence of linkage disequilibrium between closely linked loci can aid in the fine mapping of disease loci. We investigate the power of several designs for sampling individuals with different disease genotypes. As expected, haplotype data provide the greatest power for detecting disequilibrium, but, in the absence of parental information to resolve the phase of double heterozygotes, the most powerful design samples only individuals homozygous at the trait locus. For rare diseases, such a scheme is generally not feasible, and we also provide power and sample-size calculations for designs that sample heterozygotes. The results provide information useful in planning disequilibrium studies.

Chromosome Mapping↗

Multipoint linkage-disequilibrium mapping with haplotype-block structure.

The HapMap Project is providing a great deal of new information on high-resolution haplotype structure in various human populations. This information has the potential to greatly increase the power of association mapping for a fixed amount of genotyping. A number of methods have been proposed for the identification of haplotype blocks, common haplotypes, and tagging single-nucleotide polymorphisms. Here, we build on this work by developing novel methods for case-control multipoint linkage-disequilibrium (LD) mapping that gain power and speed by making explicit use of the inferred block structure. Specifically, we developed a virtual-variant approach that uses the haplotype-block information to greatly increase power for detection of untyped common variants associated with a trait. Because full multipoint LD mapping can be slow, we exploited the haplotype-block information to develop a fast single-block multipoint mapping method. Our methods are appropriate for genotype data and take into account the uncertainty in phase. We describe the methods in the context of case-parents trios, although they are also applicable to unrelated cases and controls. Our simulations indicate that the most important gains from taking into account the haplotype-block structure at the analysis stage of multipoint LD mapping come from (1) greatly increased power to detect association with untyped variants and (2) greatly improved localization of untyped variants associated with the trait. More-modest gains are obtained in improving power to detect association with a variant that is typed with a moderate amount of missing data. The methods are applied to a Crohn disease data set.

Algorithms↗

Influence of polymorphisms at loci encoding DNA repair proteins on cancer susceptibility and G2 chromosomal radiosensitivity.

Sixteen candidate polymorphisms (13 SNPs and 3 microsatellites) in nine genes from four DNA repair pathways were examined in 83 subjects, comprising 23 survivors of childhood cancer, their 23 partners, and 37 offspring, all of whom had previously been studied for G(2) chromosomal radiosensitivity. Genotype at the Asp148Glu SNP site in the APEX gene of the base excision repair (BER) pathway was associated with childhood cancer in survivors (P = 0.001, significant even after multiple test adjustment), due to the enhanced frequency of the APEX Asp148 allele among survivors in comparison to that of their partners. Analysis of variance (ANOVA) of G(2) radiosensitivity in the pooled sample, as well as family-based association test (FBAT) of the family-wise data, showed sporadic suggestions of associations between G(2) radiosensitivity and polymorphisms at two sites (the Thr241Met SNP site in the XRCC3 gene of the homologous recombinational pathway by ANOVA, and the Ser326Cys site in the hOGG1 gene of the BER pathway by FBAT analysis), but neither of these remained significant after multiple-test adjustment. This pilot study provides an intriguing indication that DNA repair gene polymorphisms may underlie cancer susceptibility and variation in radiosensitivity.

Adolescent↗

Iat epitopes on T cell receptor for self MHC class II determinants.

Anti-Iatk monoclonal antibodies (mAbs) were found to inhibit syngeneic mixed lymphocyte reaction (SMLR) of mice with k and a haplotypes (H-2k and H-2a) of the major histocompatibility complex (MHC) by acting on responder T cells but not stimulator cells. Only the early phase of SMLR was inhibited by anti-Iat mAbs. The inhibitory effect was due to the blocking of autoreactive T cells but not induction of suppressor lymphocytes. Cross-linking of Iat epitopes induced T cell proliferation of k haplotype strain. The inhibitory pattern of SMLR by four anti-Iat mAbs varied among different strains of mice. The inhibitory pattern seemed to depend on MHC and unidentified non-MHC background genes possessed by stimulator cells, suggesting that the shape of the MHC recognition site must have different conformation depending on both MHC and non-MHC gene products recognized. However, the inhibitory pattern of SMLR by anti-Iat mAbs of strains which differ at the I-J locus: B10.A(5R) and B10.A(3R) was dependent on the genotype of stimulator cells. The same was observed in B10.S(9R) and B10.HTT strains. The inhibitory activity of anti-Iat mAbs was entirely directed against responder T cells and was not passively carried out by the stimulators. It was concluded that Iat epitopes are I region controlled determinants that are utilized by T cells as receptors for self Ia antigens. It is suggested that anti-Iat mAbs react with an idiotype on the receptor(s) for Ia or I-J, or possibly a "receptor" for these receptors.

Animals↗

New immunoglobulin IgG allotypes and haplotypes found in wild mice with monoclonal anti-allotope antibodies.

Sera from 156 wild mice (Mus musculus L.) collected from parts of Eurasia, Northern Africa, and the Americas were tested in solid-phase radioimmunoassays for reactivity with 20 monoclonal anti-allotope antibodies directed against gene products of the Igh-1, Igh-3, and Igh-4 loci. Most of the wild mice have Igh phenotypes similar to those of inbred strains or heterozygotes thereof, but frequently the wild mice showed unique combinations of allotypes on a given chromosome (haplotypes) not seen in inbred strains. We found new allotypes of the Igh-1 and Igh-4 loci. Mice from three regions (Poland, Taiwan, and Japan) possess combinations of allotypes indicating that recombination, gene conversion, or other gene duplication events have occurred in portions of the Igh constant region gene complex. Most interestingly, sera of three mice from Egypt possess immunoglobulin molecules appearing to result from an intragenic recombination event involving either Igh-1d or Igh-3d with Igh-1b. Two mice from Pakistan and one from Seychelles Island in the Indian Ocean also showed similar unusual phenotypes. These mice possess two CH2 IgH-1b and two CH2 Igh-1a/CH2 Igh-1d determinants, suggesting that these variant immunoglobulin arose from recombinations within the CH2 domain.

Animals↗

Exploring genetic adaptation and microbial dynamics in engineered anaerobic ecosystems via strain-level metagenomics.

Genetic heterogeneity exists within all microbial populations, with sympatric cells of the same species often exhibiting single-nucleotide variations that influence phenotypic traits, including metabolic efficiency. However, the evolutionary dynamics of these strain-level differences in response to environmental stress remain poorly understood. Here, we present a first-of-its-kind study tracking the adaptive evolution of an anaerobic, carbon-fixing microbiota under a controlled engineered ecosystem focused on carbon dioxide bioconversion into methane. Leveraging strain-resolved metagenomics with an ad hoc variant calling and phasing approach, we mapped mutation trajectories and observed that the two dominant Methanothermobacter species maintained distinct sweeping haplotypes over time, most likely due to niche-specific metabolic roles. By combining population genetic statistics and peptide reconstruction, mer and mcrB genes emerged as potential drivers of archaeal strain-level competition. These findings pave the way for targeted engineering of microbial communities to enhance bioconversion efficiency, with significant implications for sustainable energy and carbon management in anaerobic systems.

Metagenomics↗

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article↗

Determination of haploid DNA sequences in humans: application to the glucocerebrosidase pseudogene.

Variation analyses in the human genome at the sequence level, especially human genetic population analysis and genetic epidemiology, are hampered by the difficulty to ascertain haplotypes on autosomal regions. We have designed a new methodological approach to obtain autosomal haploid sequences from diploid organisms. First, genotypes are unambiguously determined through long-range PCR and diploid DNA sequencing. Second, cloning the whole PCR-amplified segment and sequencing a single clone for those fragments that presented a heterozygous position discern the allelic phase. The second allele is deduced from the genotype, and the phase reconfirmed by sequencing a second clone. A hundred human chromosomes were analysed for a 5.4 kb encompassing the glucocerebrosidase pseudogene on human chromosome 1. Haplotypes were unambiguously ascertained for all samples. The manner to combine the used techniques makes this approach a novelty. Haploid sequences from diploid organisms are obtained in a less time consuming and more accurate manner than in other used procedures.

Glucosylceramidase↗

Phenotypic features and genetic characterization of male breast cancer families: identification of two recurrent BRCA2 mutations in north-east of Italy.

BACKGROUND: Breast cancer in men is an infrequent occurrence, accounting for approximately 1% of all breast tumors with an incidence of about 1:100,000. The relative rarity of male breast cancer (MBC) limits our understanding of the epidemiologic, genetic and clinical features of this tumor. METHODS: From 1997 to 2003, 10 MBC patients were referred to our Institute for genetic counselling and BRCA1/2 testing. Here we report on the genetic and phenotypic characterization of 10 families with MBC from the North East of Italy. In particular, we wished to assess the occurrence of specific cancer types in relatives of MBC probands in families with and without BRCA2 predisposing mutations. Moreover, families with recurrent BRCA2 mutations were also characterized by haplotype analysis using 5 BRCA2-linked dinucleotide repeat markers and 8 intragenic BRCA2 polymorphisms. RESULTS: Two pathogenic mutations in the BRCA2 gene were observed: the 9106C>T (Q2960X) and the IVS16-2A>G (splicing) mutations, each in 2 cases. A BRCA1 mutation of uncertain significance 4590C>G (P1491A) was also observed. In families with BRCA2 mutations, female breast cancer was more frequent in the first and second-degree relatives compared to the families with wild type BRCA1/2 (31.9% vs. 8.0% p = 0.001). Reconstruction of the chromosome phasing in three families and the analysis of three isolated cases with the IVS16-2A>G BRCA2 mutation identified the same haplotype associated with MBC, supporting the possibility that this founder mutation previously detected in Slovenian families is also present in the North East of our Country. Moreover, analysis of one family with the 9106C>T BRCA2 mutation allowed the identification of common haplotypes for both microsatellite and intragenic polymorphisms segregating with the mutation. Three isolated cases with the same mutation shared the same intragenic polymorphisms and three 5' microsatellite markers, but showed a different haplotype for 3' markers, which were common to all three cases. CONCLUSION: The 9106C>T and the IVS16-2A>G mutations constitute recurrent BRCA2 mutations in MBC cases from the North-East of Italy and may be associated with a founder effect. Knowledge of these two recurrent BRCA2 mutations predisposing to MBC may facilitate the analyses aimed at the identification of mutation carriers in our geographic area.

Adult↗

Computing the minimum recombinant haplotype configuration from incomplete genotype data on a pedigree by integer linear programming.

We study the problem of reconstructing haplotype configurations from genotypes on pedigree data with missing alleles under the Mendelian law of inheritance and the minimum-recombination principle, which is important for the construction of haplotype maps and genetic linkage/association analyses. Our previous results show that the problem of finding a minimum-recombinant haplotype configuration (MRHC) is in general NP-hard. This paper presents an effective integer linear programming (ILP) formulation of the MRHC problem with missing data and a branch-and-bound strategy that utilizes a partial order relationship and some other special relationships among variables to decide the branching order. Nontrivial lower and upper bounds on the optimal number of recombinants are introduced at each branching node to effectively prune the search tree. When multiple solutions exist, a best haplotype configuration is selected based on a maximum likelihood approach. The paper also shows for the first time how to incorporate marker interval distance into a rule-based haplotyping algorithm. Our results on simulated data show that the algorithm could recover haplotypes with 50 loci from a pedigree of size 29 in seconds on a Pentium IV computer. Its accuracy is more than 99.8% for data with no missing alleles and 98.3% for data with 20% missing alleles in terms of correctly recovered phase information at each marker locus. A comparison with a statistical approach SimWalk2 on simulated data shows that the ILP algorithm runs much faster than SimWalk2 and reports better or comparable haplotypes on average than the first and second runs of SimWalk2. As an application of the algorithm to real data, we present some test results on reconstructing haplotypes from a genome-scale SNP dataset consisting of 12 pedigrees that have 0.8% to 14.5% missing alleles.

Algorithms↗

Experience with cyclosporine in 1519 kidney transplantations from living donors in a national transplant programme, 1983-2002.

Following the introduction of cyclosporine as basic immunosuppression in our national transplant programme in 1983, the pool of grafts from living donors (LDs) was expanded 2 years later by also accepting LDs mismatched for 2 HLA haplotypes and living unrelated donors (LURDs), mostly spouses. A policy of approaching family members to promote donation was consistently pursued. During 1983 through 2002, nephrectomy was performed on 1519 LDs without mortality. From 1983 through 1988, our learning phase in managing cyclosporine-immunosuppression, 382 patients received first grafts from LDs. One-year graft survival (GS) rates were 94.4%, 90%, 89%, and 82% in 71 HLA identical, 260 haploidentical, 18 2-haplotypes disparate, and 33 LURD graft recipients, respectively. Corresponding half-lives were 15.8, 10.3, 11, and 9.1 years, respectively. Results improved in 1028 patients receiving first LD grafts from 1989 through 2002. Corresponding 1-year GS rates were 96.6% (n=117), 93.5% (n=650), 90.4% (n=73), and 88.8% (n=188), and half-lives were 30, 13.3, 13.5, and 12.3 years, respectively. Similar GS rates were observed in 109 recipients of repeat grafts from LDs. LDs contributed 44% and 21.6% of all first and repeat grafts transplanted, providing grafts to 11 patients (in 1983) increasing to 23 patients (in 2002) per million population per year (pmp/y). When added to grafts from cadaveric donors, 40 to 48 pmp/y were provided with a first or repeat graft since 1990, thus covering at least 65% of the national need for kidney transplantations.

Cyclosporine↗

Admixture, migrations, and dispersals in Central Asia: evidence from maternal DNA lineages.

Mitochondrial DNA (mtDNA) lineages of 232 individuals from 12 Central Asian populations were sequenced for both control region hypervariable segments, and additional informative sites in the coding region were also determined. Most of the mtDNA lineages belong to branches of the haplogroups with an eastern Eurasian (A, B, C, D, F, G, Y, and M haplogroups) or a western Eurasian (HV, JT, UK, I, W, and N haplogroups) origin, with a small fraction of Indian M lineages. This suggests that the extant genetic variation found in Central Asia is the result of admixture of already differentiated populations from eastern and western Eurasia. Nonetheless, two groups of lineages, D4c and G2a, seem to have expanded from Central Asia and might have their Y-chromosome counterpart in lineages belonging to haplotype P(xR1a). The present results suggest that the mtDNA found out of Africa might be the result of a maturation phase, presumably in the Middle East or eastern Africa, that led to haplogroups M and N, and subsequently expanded into Eurasia, yielding a geographically structured group of external branches of these two haplogroups in western and eastern Eurasia, Central Asia being a contact zone between two differentiated groups of peoples.

Africa↗

A class of population genetic questions formulated as the generalized occupancy problem.

In categorical genetic data analysis when the sampling units are classified into an arbitrary number of distinct classes, sometimes the sample size may not be large enough to apply large sample approximations for hypothesis testing purposes. Exact sampling distributions of several statistics are derived here, using combinatorial approaches parallel to the classical occupancy problem to help overcome this difficulty. Since the multinomial probabilities can be unequal, this situation is described as a generalized occupancy problem. The sampling properties derived are used to examine nonrandomness of occurrence of mutagen-induced mutations across loci, to devise tests of Hardy-Weinberg proportions of genotype frequencies in the presence of a large number of alleles, and to provide a global test of gametic phase disequilibrium of several restriction site polymorphisms.

Genetics, Population↗

Beta S haplotypes in various world populations.

We have determined the beta S haplotypes in 709 patients with sickle cell anemia, 30 with SC disease, 91 with S-beta-thalassemia, and in 322 Hb S heterozygotes from different countries. The methodology concerned the detection of mutations in the promoter sequences of the G gamma- and A gamma-globin genes through dot blot analysis of amplified DNA with 32P-labeled probes, and an analysis of isolated Hb F by reversed phase high performance liquid chromatography to detect the presence of the A gamma T chain [A gamma 75(E19)Ile----Thr] that is characteristic for haplotype 17 (Cameroon). The results support previously published data obtained with conventional methodology that indicates that the beta S gene arose separately in different locations. The present methodology has the advantage of being relatively inexpensive and fast, allowing the collection of a vast body of data in a short period of time. It also offers the opportunity of identifying unusual beta S haplotypes that may be associated with a milder expression of the disease. The numerous blood samples obtained from many SS patients living in different countries made it possible to compare their hematological data. Such information is included (as average values) for 395 SS patients with haplotype 19/19, for 2 with haplotype 17/17, for 50 with haplotype 20/20, for 2 with haplotype 3/3, and for 37 with haplotype 31/31. Some information on haplotype characteristics of normal beta A chromosomes is also presented.

Africa↗

Genetic analysis of case/control data using estimated haplotype frequencies: application to APOE locus variation and Alzheimer's disease.

There is growing debate over the utility of multiple locus association analyses in the identification of genomic regions harboring sequence variants that influence common complex traits such as hypertension and diabetes. Much of this debate concerns the manner in which one can use the genotypic information from individuals gathered in simple sampling frameworks, such as the case/control designs, to actually assess the association between alleles in a particular genomic region and a trait. In this paper we describe methods for testing associations between estimated haplotype frequencies derived from multilocus genotype data and disease endpoints assuming a simple case/control sampling design. These proposed methods overcome the lack of phase information usually associated with samples of unrelated individuals and provide a comprehensive way of assessing the relationship between sequence or multiple-site variation and traits and diseases within populations. We applied the proposed methods in a study of the relationship between polymorphisms within the APOE gene region and Alzheimer's disease. Cases and controls for this study were collected from the United States and France. Our results confirm the known association between the APOE locus and Alzheimer's disease, even when the epsilon 4 polymorphism is not contained in the tested haplotypes. This suggests that, in certain situations, haplotype information and linkage disequilibrium-induced associations between polymorphic loci that neighbor loci harboring functional sequence variants can be exploited to identify disease-predisposing alleles in large, freely mixing populations via estimated haplotype frequency methods.

Aged↗