Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Testicular carcinoma and HLA Class II genes.

BACKGROUND: The association with histocompatibility antigens (HLA), in particular Class II genes (DQB1, DRB1), has recently been suggested to be one of the genetic factors involved in testicular germ cell tumor (TGCT) development. The current study, which uses genotyping of microsatellite markers, was designed to replicate previous associations. METHODS: In 151 patients, along with controls comprising parents or spouses, the HLA region (particularly Class II) on chromosome 6p21 was genotyped for a set of 15 closely linked microsatellite markers. RESULTS: In both patients and controls, strong linkage disequilibrium was observed in the genotyped region, indicating that similar haplotypes are likely to be identical by descent. However, association analysis and the transmission disequilibrium test did not show significant results. Haplotype sharing statistics, a haplotype method that derives extra information from phase and single marker tests, did not show differences in haplotype sharing between patients and controls. CONCLUSION: The current genotyping study did not confirm the previously reported association between HLA Class II genes and TGCT. As the HLA alleles for which associations were reported are also prevalent in the Dutch populations, these associations are likely to be nonexistent or much weaker than previously reported.

Adolescent↗

Angiotensin I-converting enzyme (ACE): estimation of DNA haplotypes in unrelated individuals using denaturing gradient gel blots.

The angiotensin I-converting enzyme (ACE) gene (17q23) is a candidate gene for essential hypertension and related diseases, but investigation of its role in human pathology is hampered by a lack of identified polymorphisms. Currently, a 287-bp insertion/deletion (I/D) RFLP in intron 16 represents the only one known. Additional polymorphisms for the ACE gene would make most families informative for linkage studies and would allow haplotypes to be assigned in association studies. To increase the information provided by the ACE gene, we used a sensitive screening technique, denaturing gradient gel electrophoresis (DGGE) blots, to identify polymorphisms and combined this with gene counting to identify haplotypes. Five independent polymorphisms, restriction fragment melting polymorphisms (RFMPs), were identified by four probes (encompassing half of the ACE cDNA) in digests produced by three restriction enzymes (DdeI, RsaI, and AluI). One RFMP has three alleles while the others have two alleles. In a sample of 67 unrelated control subjects, minor allele frequencies ranged from 0.12 to 0.49. A significant level of linkage disequilibrium was found for all pairs of markers. The four most informative RFMPs, taken in combination, define 24 potential haplotypes. Based on gene counting, 11 of the 24 are rare or nonexistent in this population, and the estimated heterozygosity of the remaining 13 haplotypes approaches 80%. Under these conditions for the ACE locus, phase-unknown genotypes could be assigned to haplotype pairs in unrelated subjects with reasonable certainty. Thus, using DGGE blot technique for identifying numerous DNA polymorphisms in a candidate locus, in combination with gene counting, one can often identify DNA haplotypes for both related and unrelated study subjects at a candidate locus. These markers in the ACE gene should be useful for clinical and epidemiologic studies of the role of ACE in human disease.

Alleles↗

Association of specific interleukin 1 gene cluster polymorphisms with increased susceptibility for Behcet's disease.

OBJECTIVE: The aim of this study was to investigate if the inheritance of specific polymorphisms of interleukin 1 (IL-1) A, IL-1B and IL-1 receptor antagonist (IL-1RN) genes could affect the susceptibility to Behçet's disease (BD). METHODS: A total of 132 BD patients and 105 healthy controls were genotyped for IL-1A -889, IL-1B -511, -35, +5810, +5887, and IL-1RN +8006, +8061, +9589, +11,100 single nucleotide polymorphisms, and IL-1RN 86-bp variable number of tandem repeat polymorphism. chi 2-analysis was used to compare the allele and genotype frequencies of the cases and controls. IL-1A and IL-1B haplotypes were reconstructed using the Phase program. RESULTS: Inheritance of the C allele of the IL-1A -889 polymorphism was associated with BD (OR=2.0, P=0.01) and inheritance of the IL-1A -889C/IL-1B +5887T haplotype was identified as an increased risk for BD. The IL-1A -889 and IL-1B +5887 CC/TT combined genotype was significantly more observed in BD cases than in controls (57.5 vs 38.1%, OR=2.2, P=0.003). No association with BD was found for other investigated polymorphisms in the IL-1B and IL-1RN genes. CONCLUSION: Susceptibility to BD is increased in individuals carrying both the IL-1A -889C and IL-1B +5887T haplotype. Individuals who are both homozygous CC at IL-1A -889 and TT at IL-1B +5887 appear to have twice the risk of developing BD as individuals having other IL-1A -889/IL-1B +5887 genotypes.

Behcet Syndrome↗

Estimating haplotype-disease associations with pooled genotype data.

The genetic dissection of complex human diseases requires large-scale association studies which explore the population associations between genetic variants and disease phenotypes. DNA pooling can substantially reduce the cost of genotyping assays in these studies, and thus enables one to examine a large number of genetic variants on a large number of subjects. The availability of pooled genotype data instead of individual data poses considerable challenges in the statistical inference, especially in the haplotype-based analysis because of increased phase uncertainty. Here we present a general likelihood-based approach to making inferences about haplotype-disease associations based on possibly pooled DNA data. We consider cohort and case-control studies of unrelated subjects, and allow arbitrary and unequal pool sizes. The phenotype can be discrete or continuous, univariate or multivariate. The effects of haplotypes on disease phenotypes are formulated through flexible regression models, which allow a variety of genetic hypotheses and gene-environment interactions. We construct appropriate likelihood functions for various designs and phenotypes, accommodating Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We develop simple and efficient numerical algorithms for calculating the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. We assess the performance of the proposed methods through simulation studies, and provide an application to the Finland-United States Investigation of NIDDM Genetics Study. The results show that DNA pooling is highly efficient in studying haplotype-disease associations. As a by-product, this work provides valid and efficient methods for estimating haplotype-disease associations with unpooled DNA samples.

Algorithms↗

Direct molecular haplotyping of long-range genomic DNA with M1-PCR.

Haplotypes, combinations of several phase-determined polymorphic markers, are extremely valuable for studies of disease association and chromosome evolution. Here we describe a technique called M1-PCR (M for "multiplex" and 1 for "single-copy DNA molecules") that enables direct molecular haplotyping of several polymorphic markers separated by as many as 24 kb. A genomic DNA sample first is diluted to approximately single-copy. The haplotype is directly determined by simultaneously genotyping several polymorphic markers in the same reaction with a multiplex PCR and base extension reaction. This approach does not rely on pedigree data and does not require previous amplification of the entire genomic region containing the selected markers.

Carrier Proteins↗

VEGF, FGF1, FGF2 and EGF gene polymorphisms and psoriatic arthritis.

BACKGROUND: Angiogenesis appears to be a first-order event in psoriatic arthritis (PsA). Among angiogenic factors, the cytokines vascular endothelial growth factor (VEGF), epidermal growth factor (EGF), and fibroblast growth factors 1 and 2 (FGF1 and FGF2) play a central role in the initiation of angiogenesis. Most of these cytokines have been shown to be upregulated in or associated with psoriasis, rheumatoid arthritis (RA) or ankylosing spondylitis (AS). As these diseases share common susceptibility associations with PsA, investigation of these angiogenic factors is warranted. METHODS: Two hundred and fifty-eight patients with PsA and 154 ethnically matched controls were genotyped using a Sequenom chip-based MALDI-TOF mass spectrometry platform. Four SNPs in the VEGF gene, three SNPs in the EGF gene and one SNP each in FGF1 and FGF2 genes were evaluated. Statistical analysis was performed using Fisher's exact test, and the Cochrane-Armitage trend test. Associations with haplotypes were estimated by using weighted logistic models, where the individual haplotype estimates were obtained using Phase v2.1. RESULTS: We have observed an increased frequency in the T allele of VEGF +936 (rs3025039) in control subjects when compared to our PsA patients [Fisher's exact p-value = 0.042; OR 0.653 (95% CI: 0.434, 0.982)]. Haplotyping of markers revealed no significant associations. CONCLUSION: The T allele of VEGF in +936 may act as a protective allele in the development of PsA. Further studies regarding the role of pro-angiogenic markers in PsA are warranted.

Adult↗

An approximation algorithm for haplotype inference by maximum parsimony.

This paper studies haplotype inference by maximum parsimony using population data. We define the optimal haplotype inference (OHI) problem as given a set of genotypes and a set of related haplotypes, find a minimum subset of haplotypes that can resolve all the genotypes. We prove that OHI is NP-hard and can be formulated as an integer quadratic programming (IQP) problem. To solve the IQP problem, we propose an iterative semidefinite programming-based approximation algorithm, (called SDPHapInfer). We show that this algorithm finds a solution within a factor of O(log n) of the optimal solution, where n is the number of genotypes. This algorithm has been implemented and tested on a variety of simulated and biological data. In comparison with three other methods, (1) HAPAR, which was implemented based on the branching and bound algorithm, (2) HAPLOTYPER, which was implemented based on the expectation-maximization algorithm, and (3) PHASE, which combined the Gibbs sampling algorithm with an approximate coalescent prior, the experimental results indicate that SDPHapInfer and HAPLOTYPER have similar error rates. In addition, the results generated by PHASE have lower error rates on some data but higher error rates on others. The error rates of HAPAR are higher than the others on biological data. In terms of efficiency, SDPHapInfer, HAPLOTYPER, and PHASE output a solution in a stable and consistent way, and they run much faster than HAPAR when the number of genotypes becomes large.

Algorithms↗

Complex haplotypes of IRS2 gene are associated with severe obesity and reveal heterogeneity in the effect of Gly1057Asp mutation.

In order to understand the role of the insulin receptor substrate-2 (IRS2) gene (chromosome region: 13q34) in obesity, a complex disorder associated with insulin resistance and glucose intolerance, we determined single nucleotide polymorphims (SNPs) and complex haplotypes in women with morbid obesity and a body mass index (BMI) of 41+/-0.8 kg/m2 ( n=99) compared with controls having a BMI of 23.8+/-0.1 kg/m2 ( n=92). Sequencing of unphased DNA or digestion of polymerase chain reaction fragments revealed seven SNPs, including a new C/T(-769) replacement at the 5' untranslated region. Considering four or seven SNPs, we reconstructed with the PHASE program nine or 24 haplotypes, respectively, that were well correlated into the cladogram. Logistic regression analysis with nine haplotypes in the whole sample revealed that obesity was associated with haplotype H3, with P<0.025, an odds ratio (OR) of 1.9 and a 95% confidence interval (CI) of 1.1-3.4, or pairs 3/3 ( P<0.005, OR=8.7, CI=1.9-40.1) and 3/4 ( P<0.023, OR=2.5, CI=1.1-5.6), all containing the the Gly1057Asp allelic variant of IRS2, whereas controls were associated with H5 and H6 ( P<0.02, OR=0.2, CI=0.01-0.81). Although obese H5 carriers (also containing Gly1057Asp mutation) were the most insulin resistant, haplotypes of IRS2 were poorly correlated (analysis of variance) with insulin resistance. By contrast, haplotypes H3, H4 and pairs 3/3 were consistently associated with increased 2-h glucose levels during an oral glucose tolerance test in obese individuals ( P<0.0005, 0.025 and 0.027, respectively). These data indicate that IRS2 is an influential gene in severe obesity and glucose intolerance in this population, whereas gene-based haplotypes of IRS2 have revealed heterogeneity in the behaviour of the Gly1057Asp mutation in relation to insulin resistance.

Adult↗

SNPs, haplotypes, and model selection in a candidate gene region: the SIMPle analysis for multilocus data.

Modern molecular techniques make discovery of numerous single nucleotide polymorphims (SNPs) in candidate gene regions feasible. Conventional analysis relies on either independent tests with each variant or the use of haplotypes in association analysis. The first technique ignores the dependencies between SNPs. The second, though it may increase power, often introduces uncertainty by estimating haplotypes from population data. Additionally, as the number of loci expands for a haplotype, ambiguity in interpretation increases for determining the underlying genetic components driving a detected association. Here, we present a genotype-level analysis to jointly model the SNPs via a SNP interaction model with phase information (SIMPle) to capture the underlying haplotype structure. This analysis estimates both the risk associated with each variant and the importance of phase between pairwise combinations of SNPs. Thus, rather than selecting between genotype- or haplotype-level approaches, the SIMPle method frames the analysis of multilocus data in a model selection paradigm, the aim to determine which SNPs, phase terms, and linear combinations best describe the relation between genetic variation and a trait of interest. To avoid unstable estimation due to sparse data and to incorporate both the dependencies among terms and the uncertainty in model selection, we propose a Bayes model averaging procedure. This highlights key SNPs and phase terms and yields a set of best representative models. Using simulations, we demonstrate the utility of the SIMPle model to identify crucial SNPs and underlying haplotype structures across a variety of causal models and genetic architectures.

Bayes Theorem↗

Ancient origin of the CAG expansion causing Huntington disease in a Spanish population.

Huntington disease (HD) is an autosomal dominant neurodegenerative disorder characterized clinically by progressive motor impairment, cognitive decline, and emotional deterioration. The disease is caused by the abnormal expansion of a CAG trinucleotide repeat in the first exon of the huntingtin gene in chromosome 4p16.3. HD is spread worldwide and it is generally accepted that few mutational events account for the origin of the pathogenic CAG expansion in most populations. We have investigated the genetic history of HD mutation in 83 family probands from the Land of Valencia, in Eastern Spain. An analysis of the HD/CCG repeat in informative families suggested that at least two main chromosomes were associated in the Valencian population, one associated with allele 7 (77 mutant chromosomes) and one associated with allele 10 (two mutant chromosomes). Haplotype A-7-A (H1) was observed in 47 out of 48 phase-known mutant chromosomes, obtained by segregation analysis, through the haplotype analysis of rs1313770-HD/CCG-rs82334, as it also was in 120 out of 166 chromosomes constructed by means of the PHASE program. The genetic history and geographical distribution of the main haplotype H1 were both studied by constructing extended haplotypes with flanking short tandem repeats (STRs) D4S106 and D4S3034. We found that we were able to determine the age of the CAG expansion associated with the haplotype H1 as being between 4,700 and 10,000 years ago. Furthermore, we observed a nonhomogenous distribution in the different regions associated with the different extended haplotypes of the ancestral haplotype H1, suggesting that local founder effects have occurred.

Alleles↗

Haplotype and missing data inference in nuclear families.

Determining linkage phase from population samples with statistical methods is accurate only within regions of high linkage disequilibrium (LD). Yet, affected individuals in a genetic mapping study, including those involving cases and controls, may share sequences identical-by-descent stretching on the order of 10s to 100s of kilobases, quite possibly over regions of low LD in the population. At the same time, inferring phase from nuclear families may be hampered by missing family members, missing genotypes, and the noninformativity of certain genotype patterns. In this study, we reformulate our previous haplotype reconstruction algorithm, and its associated computer program, to phase parents with information derived from population samples as well as from their offspring. In applications of our algorithm to 100-kb stretches, simulated in accordance to a Wright-Fisher model with typical levels of LD in humans, we find that phase reconstruction for 160 trios with 10% missing data is highly accurate (>90%) over the entire length. Furthermore, our algorithm can estimate allelic status for missing data at high accuracy (>95%). Finally, the input capacity of the program is vast, easily handling thousands of segregating sites in > or = 1000 chromosomes.

Algorithms↗

The polymorphism and haplotypes of XRCC1 and survival of non-small-cell lung cancer after radiotherapy.

PURPOSE: The X-ray repair cross-complementing Group 1 (XRCC1) protein is involved mainly in the base excision repair of the DNA repair process. This study examined the association of 3 polymorphisms (codon 194, 280, and 399) of XRCC1 and lung cancer in terms of whether or not these polymorphisms have an effect on the survival of lung cancer patients who have received radiotherapy. METHODS AND MATERIALS: Between January 2000 and April 2004, 229 lung cancer patients with non-small-cell lung cancer in Stages I-III were enrolled. Genotyping was performed by single base primer extension assay using the SNP-IT Kit with genomic DNA samples from all patients. The haplotype of the XRCC1 polymorphisms was estimated by PHASE version 2.1. RESULTS: The patients consisted of 191 (83.4%) males and 38 (16.6%) females with a median age of 62 (range, 26-88 years). Sixty percent of the patients were included in Stage I-IIIa. The median progression-free and overall survival was 13 months and 16 months, respectively. The XRCC1 codon 194, histology, and stage were shown to be significant predictors of the progression-free survival. The 6 haplotypes among the XRCC1 polymorphisms (194, 280, and 399) were estimated by PHASE v.2.1. The patients with haplotype pairs other than the homozygous TGG haplotype pairs survived significantly longer (p = 0.04). CONCLUSIONS: Polymorphisms of XRCC1 have an effect on the survival of lung cancer patients treated with radiotherapy, and this effect seems to be more significant after the haplotype pairs are considered.

Adult↗

Low-activity haplotype of the microsomal epoxide hydrolase gene is protective against placental abruption.

OBJECTIVE: We wanted to determine whether genetic variability in the gene encoding microsomal epoxide hydrolase (EPHX) contributes to individual differences in susceptibility to the occurrence of placental abruption. METHODS: The study involved 117 women with placental abruption and 115 healthy control pregnant women who were genotyped for two single nucleotide polymorphisms (SNPs), T-->C (Tyr113His) in exon 3 and A-->G (His139Arg) in exon 4, in the EPHX gene. Chi-square analysis was used to assess genotype and allele frequency differences between the women with placental abruption and the control group. In addition, single-point analysis was expanded to pair of loci haplotype analysis to examine the estimated haplotype frequencies of the two SNPs, of unknown phase, among the women with placental abruption and the control group. Estimated haplotype frequencies were assessed using the maximum-likelihood method, employing an expectation-maximization algorithm. RESULTS: Single-point allele and genotype distributions in exons 3 and 4 of the EPHX gene were not statistically different between the groups. However, in the haplotype estimation analysis we observed a significantly decreased frequency of haplotype C-A (His113-His139) among the placental abruption group compared with the control group (P = .007). The odds ratio for placental abruption associated with the low-activity haplotype C-A (His113-His139) was 0.552 (95% confidence interval, 0.358 to 0.851). CONCLUSIONS: The use of two intragenic SNPs jointly in haplotype analysis of association demonstrated that the genetically determined low-activity haplotype C-A (His113-His139) was significantly less frequent in women with placental abruption.

Abruptio Placentae↗

A polymorphism in the promoter region of the CD86 (B7.2) gene is associated with systemic sclerosis.

Systemic sclerosis (SSc) is a connective tissue disease of unknown aetiology characterized by fibrosis of the skin and internal organs, vascular abnormalities and humoral autoimmunity. Strong T-cell-dependent autoantibody and HLA associations are found in SSc subsets. The co-stimulatory molecule, CD86, expressed by antigen-presenting cells, plays a crucial role in priming naïve lymphocytes. We hypothesized that SSc, or one of the disease subsets, could be associated with single-nucleotide polymorphisms of the CD86 gene. Using sequence specific primer-polymerase chain reaction (SSP-PCR) methodology, we assessed four CD86 polymorphisms in 221 patients with SSc and 227 healthy control subjects from the UK. Haplotypes were constructed by inference and confirmed using PHASE algorithm. We found a strong association between SSc and a specific haplotype (haplotype 5), which was more prevalent in patients than in controls (29% vs 15%, OR = 2.3, chi(2) = 12, P = 0.0005). This association could be attributed to the novel -3479 promoter polymorphism; a significant difference was observed in the distribution of the CD86 -3479 G allele in patients with SSc compared to controls (43.7% vs. 32.4%, OR = 1.7, chi(2) = 12.1, P = 0.0005). TRANSFAC analyses suggest that the CD86-3479T allele contains putative GATA and TBP sites, whereas G allele does not. We assessed the relative DNA protein-binding activity of the -3479 polymorphism in vitro using electromobility gel shift assays (EMSA), which showed that the -3479G allele has less binding affinity compared to the T allele for nuclear proteins. These findings highlight the importance of co-stimulatory pathways in SSc pathogenesis.

Algorithms↗

Risk of trachomatous scarring and trichiasis in Gambians varies with SNP haplotypes at the interferon-gamma and interleukin-10 loci.

Experimental evidence implicates interferon gamma (IFNgamma) in protection from and resolution of chlamydial infection. Conversely, interleukin 10 (IL10) is associated with susceptibility and persistence of infection and pathology. We studied genetic variation within the IL10 and IFNgamma loci in relation to the risk of developing severe complications of human ocular Chlamydia trachomatis infection. A total of 651 Gambian subjects with scarring trachoma, of whom 307 also had potentially blinding trichiasis and pair-matched controls with normal eyelids, were screened for associations between single-nucleotide polymorphisms (SNPs), SNP haplotypes and the risk of disease. MassEXTEND (Sequenom) and MALDI-TOF mass spectrometry were used for detection and analysis of SNPs and the programs PHASE and SNPHAP used to infer haplotypes from population genetic data. Multivariate conditional logistic regression analysis identified IL10 and IFNgamma SNP haplotypes associated with increased risk of both trachomatous scarring and trichiasis. SNPs in putative IFNgamma and IL10 regulatory regions lay within the disease-associated haplotypes. The IFNgamma +874A allele, previously linked to lower IFNgamma production, lies in the IFNgamma risk haplotype and was more common among cases than controls, but not significantly so. The promoter IL10-1082G allele, previously associated with high IL10 expression, is in both susceptibility and resistance haplotypes.

Alleles↗

Notes on the maximum likelihood estimation of haplotype frequencies.

The maximum likelihood estimation (MLE) is one of the most popular ways to estimate haplotype frequencies of a population with genotype data whose linkage phases are unknown. The MLE is commonly implemented in the use of the Expectation-Maximization (EM) algorithm. It is known that the EM algorithm carries the risk that an estimator may converge erroneously to one of the local maxima or saddle points of the likelihood surface, resulting in serious errors in the MLE of haplotype frequencies. In this note, by theoretical treatments we present the necessary and sufficient conditions that the local maxima or saddle points on the likelihood surface appear. As a rule of thumb, that the difference between the coupling and repulsive haplotype frequencies in phase known individuals is 3/2 times larger than the frequency of phase ambiguous individuals is the sufficient condition that the likelihood surface is unimodal. Moreover, we present the analytic solution to the biallelic two-locus problem, and construct a general algorithm to obtain the global maximum.

Algorithms↗

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics↗

Defining the contribution of the HLA region to cis DQ2-positive coeliac disease patients.

The major genetic susceptibility to coeliac disease is contributed by the human leukocyte antigen (HLA) region. The primary association is with the HLA-DQ2 molecule, encoded by the DQA1*05 and DQB1*02 alleles, which is expressed by over 90% of patients. The aim of our study was to perform an extensive scan of the entire HLA region to determine whether there is evidence for the presence of additional HLA susceptibility genes for coeliac disease in the Dutch population, acting independently of DQ2. In all, 16 microsatellite markers and the DQA1 and DQB1 genes were genotyped in simplex cis DQ2-positive coeliac disease families and cis DQ2-positive control families. Allele frequencies of markers on phase-known DQ2-positive haplotypes transmitted to patients were compared to a combined group of DQ2-positive nontransmitted and control haplotypes, thereby controlling for the DQ2 contribution. No significant differences at any of the marker loci were detected, suggesting that DQ2 is the major HLA risk factor for coeliac disease. Individuals homozygous for DQ2 or heterozygous for DQA1*05-DQB1*02/DQA1*0201-DQB1*02 were found to be at five-fold increased risk for development of coeliac disease (P<10(-8)). This risk seems to be conferred by the presence of a second DQB1*02 allele next to one DQA1*05-DQB1*02 haplotype, independently of the second DQA1 allele.

Adolescent↗