Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

The effect of HLA-DR on susceptibility to rheumatoid arthritis is influenced by the associated lymphotoxin alpha-tumor necrosis factor haplotype.

OBJECTIVE: HLA-DRB1, a major genetic determinant of susceptibility to rheumatoid arthritis (RA), is located within 1,000 kb of the gene encoding tumor necrosis factor (TNF). Because certain HLA-DRB1*04 subtypes increase susceptibility to RA, investigation of the role of the TNF gene is complicated by linkage disequilibrium (LD) between TNF and DRB1 alleles. By adequately controlling for this LD, we aimed to investigate the presence of additional major histocompatibility complex (MHC) susceptibility genes. METHODS: We identified 274 HLA-DRB1*04-positive cases of RA and 271 HLA-DRB1*04-positive population controls. Each subject was typed for 6 single-nucleotide polymorphisms within a 4.5-kb region encompassing TNF and lymphotoxin alpha (LTA). LTA-TNF haplotypes in these unrelated individuals were determined using a combination of family data and the PHASE software program. RESULTS: Significant differences in LTA-TNF haplotype frequencies were observed between different subtypes of HLA-DRB1*04. The LTA-TNF haplotypes observed were very restricted, with only 4 haplotypes constituting 81% of all haplotypes present. Among individuals carrying DRB1*0401, the LTA-TNF 2 haplotype was significantly underrepresented in cases compared with controls (odds ratio 0.5 [95% confidence interval 0.3-0.8], P = 0.007), while in those with DRB1*0404, the opposite effect was observed (P = 0.007). CONCLUSION: These findings suggest that the MHC contains genetic elements outside the LTA-TNF region that modify the effect of HLA-DRB1 on susceptibility to RA.

Arthritis, Rheumatoid↗

Maximum-likelihood estimation of haplotype frequencies in nuclear families.

The importance of haplotype analysis in the context of association fine mapping of disease genes has grown steadily over the last years. Since experimental methods to determine haplotypes on a large scale are not available, phase has to be inferred statistically. For individual genotype data, several reconstruction techniques and many implementations of the expectation-maximization (EM) algorithm for haplotype frequency estimation exist. Recent research work has shown that incorporating available genotype information of related individuals largely increases the precision of haplotype frequency estimates. We, therefore, implemented a highly flexible program written in C, called FAMHAP, which calculates maximum likelihood estimates (MLEs) of haplotype frequencies from general nuclear families with an arbitrary number of children via the EM-algorithm for up to 20 SNPs. For more loci, we have implemented a locus-iterative mode of the EM-algorithm, which gives reliable approximations of the MLEs for up to 63 SNP loci, or less when multi-allelic markers are incorporated into the analysis. Missing genotypes can be handled as well. The program is able to distinguish cases (haplotypes transmitted to the first affected child of a family) from pseudo-controls (non-transmitted haplotypes with respect to the child). We tested the performance of FAMHAP and the accuracy of the obtained haplotype frequencies on a variety of simulated data sets. The implementation proved to work well when many markers were considered and no significant differences between the estimates obtained with the usual EM-algorithm and those obtained in its locus-iterative mode were observed. We conclude from the simulations that the accuracy of haplotype frequency estimation and reconstruction in nuclear families is very reliable in general and robust against missing genotypes.

Algorithms↗

History of Lipizzan horse maternal lines as revealed by mtDNA analysis.

Sequencing of the mtDNA control region (385 or 695 bp) of 212 Lipizzans from eight studs revealed 37 haplotypes. Distribution of haplotypes among studs was biased, including many private haplotypes but only one haplotype was present in all the studs. According to historical data, numerous Lipizzan maternal lines originating from founder mares of different breeds have been established during the breed's history, so the broad genetic base of the Lipizzan maternal lines was expected. A comparison of Lipizzan sequences with 136 sequences of domestic- and wild-horses from GenBank showed a clustering of Lipizzan haplotypes in the majority of haplotype subgroups present in other domestic horses. We assume that haplotypes identical to haplotypes of early domesticated horses can be found in several Lipizzan maternal lines as well as in other breeds. Therefore, domestic horses could arise either from a single large population or from several populations provided there were strong migrations during the early phase after domestication. A comparison of Lipizzan haplotypes with 56 maternal lines (according to the pedigrees) showed a disagreement of biological parentage with pedigree data for at least 11% of the Lipizzans. A distribution of haplotype-frequencies was unequal (0.2%-26%), mainly due to pedigree errors and haplotype sharing among founder mares.

Animals↗

Finding haplotype block boundaries by using the minimum-description-length principle.

We present a method for detecting haplotype blocks that simultaneously uses information about linkage-disequilibrium decay between the blocks and the diversity of haplotypes within the blocks. By use of phased single-nucleotide polymorphism data, our method partitions a chromosome into a series of adjacent, nonoverlapping blocks. The partition is made by choosing among a family of Markov models for block structure in a chromosomal region. Specifically, in the model, the occurrence of haplotypes within blocks follows a time-inhomogeneous Markov process along the chromosome, and we choose among possible partitions by using the two-stage minimum-description-length criterion. When applied to data simulated from the coalescent with recombination hotspots, our method reliably situates block boundaries at the hotspots and infrequently places block boundaries at sites with background levels of recombination. We apply three previously published block-finding methods to the same data, showing that they either are relatively insensitive to recombination hotspots or fail to discriminate between background sites of recombination and hotspots. When applied to the 5q31 data of Daly et al., our method identifies more block boundaries in agreement with those found by Daly et al. than do other methods. These results suggest that our method may be useful for designing association-based mapping studies that exploit haplotype blocks.

Algorithms↗

Single-molecule dilution and multiple displacement amplification for molecular haplotyping.

Separate haploid analysis is frequently required for heterozygous genotyping to resolve phase ambiguity or confirm allelic sequence. We demonstrate a technique of single-molecule dilution followed by multiple strand displacement amplification to haplotype polymorphic alleles. Dilution of DNA to haploid equivalency, or a single molecule, is a simple method for separating di-allelic DNA. Strand displacement amplification is a robust method for non-specific DNA expansion that employs random hexamers and phage polymerase Phi29 for double-stranded DNA displacement and primer extension, resulting in high processivity and exceptional product length. Single-molecule dilution was followed by strand displacement amplification to expand separated alleles to microgram quantities of DNA for more efficient haplotype analysis of heterozygous genes.

Alleles↗

Influence of the HLA-DR4 antigen and iodine status on the development of autoimmune postpartum thyroiditis.

HLA-A, -B, and -DR antigens were determined in all 50 women with a serum thyroid microsomal hemagglutination antibody (MsAb) titer equal to or greater than 1:100 in the first trimester of pregnancy in a population of 733 pregnant women. The DR4 antigen was found in 58.0% of the women compared to 33.7% in control subjects, which corresponds to a relative risk of 2.71 (P less than 0.01 by X2 test). The MsAb-positive women were examined regularly during the year after delivery for the development of thyroid dysfunction. The DR4 antigen frequency was found to be even higher, 69.0% (relative risk = 4.36; P less than 0.001), among the 29 women who developed hypothyroidism in the postpartum period. No other HLA antigen deviations were found among those 15 hypothyroid women in whom an initial thyrotoxic phase occurred before hypothyroidism. The B8, DR3 haplotype was found in 3 of 5 women who developed Graves' thyrotoxicosis alone. Urinary iodine excretion measured in some MsAb-positive women 3 (n = 19) or 6 months (n = 29) postpartum, respectively, was compatible with leakage of thyroid iodine during the initial destruction-induced thyrotoxic phase of postpartum thyroiditis, followed by low iodine excretion during the subsequent hypothyroid phase. We conclude that genes coding for the DR4 antigen may have a regulatory influence on MsAb production, which in turn affects the development of postpartum hypothyroidism. Thyroid iodine content and iodine intake also may have an impact on the severity of the thyrotoxic and the hypothyroid phases of autoimmune postpartum thyroiditis.

Adult↗

Host genetics and resistance to acute Trypanosoma cruzi infection in mice. I. Antibody isotype profiles.

Some strains of inbred mice survive acute infection with Trypanosoma cruzi while others die within a few weeks after infection. Mice which express B10 background genes and either the H-2q or H-2d haplotypes are resistant and survive. However, mice which share the B10 genetic background but express H-2k alleles die, usually within 4 weeks following infection. These data confirm that at least 1 gene in the major histocompatibility complex can determine whether an animal lives or dies during the acute phase. Expression of the H-2q haplotype on the B10 genetic background or in DBA/1 mice is associated with resistance, but H-2q mice expressing the C3H background are susceptible. Therefore, at least 1 gene in the genetic background also influences resistance. Our data suggest that genes associated with resistance must be present in both the MHC and the genetic background or the animal will die. The isotypes and specificities of parasite reactive antibodies found in the serum of different inbred mouse strains were assessed during acute infection. Levels of IgM were higher in sera from mice which express the resistant B10 background than in sera from mice expressing the susceptible C3H background. Conversely, mice which share the C3H background genes produced high levels of anti-parasite IgG2a when compared to B10 congenic strains. Antigen specificity, however, may be influenced by both background and MHC genes, as congenic strains expressing different MHC haplotypes recognized different constellations of T. cruzi antigens.(ABSTRACT TRUNCATED AT 250 WORDS)

Acute Disease↗

[Single nucleotide polymorphism and haplotype in TBX1 gene of patients with conotruncal defects: analysis of 130 cases].

OBJECTIVE: To investigate the distribution of the single nucleotide polymorphism (SNP) sites in TBX1 gene and the distribution of related haplotypes in the patients with conotruncal defects (CTD) and normal people. METHODS: The genotypes of the 3 selected SNPs: G2857C (rs737868), G2963A (rs28649236), and A6571T (rs28939675) in TBX1 gene were analyzed by PCR-RFLP among 130 patients with CTD and 200 normal people. Contingency table was applied to analyze the frequencies of these SNP genotypes and related alleles. PHASE software was used to construct the haplotypes and analyze the haplotype frequencies in these 2 groups. RESULTS: There were no significant differences in the allele frequency and genotype rates of the SNPs G2587C and A6571T between the CTD patients and normal controls (all P > 0.05). However, the allele frequency and genotype rates of the SNP G2963A were significant different between he CTD patients and normal controls: the G allele frequency in the CTD patients was 53.8%, significantly higher than that in the normal controls (42.5%, chi(2) = 8.14, P < 0.005); and the AA genotype rate of the CTD patients was 21.6%, significantly lower than that of the controls (38.0%), and the GA genotype rate in the CTD patients was 49.2%, significantly higher than that in the controls (39.0%) (both chi(2) = 9.9, P < 0.05). The haplotype frequencies of G2587/G2963/A6571 and G2587/A2963/T6571 of the CTD patients were 49.2% and 14.6% respectively, both significantly higher than those of the normal controls (36.3% and 9.5% respectively), and the haplotype frequencies of G2587/G2963/T6571 and G2587/A2963/A6571 in the CTD patients were 34.6% and 3% respectively, both significantly lower than those in the normal controls (48.3% and 18% respectively) (chi(2) = 22.39, P < 0.005). CONCLUSION: The SNP site G2963A located in the coding-region of TBX1 gene is associated with CTD. The persons with G2963 have higher risk of CTD than those with A2963. The haplotypes constructed with these 3 SNP sites may be linked with the susceptibility gene of CTD.

Adolescent↗

The locus ordering problem.

Studies of phenotypes defined by codominant alleles at two or more loci in three-generation families allow haplotypes to be deduced. These data are easily summarized by the Mendelian convention of upper and lower case, with case defining phase rather than nature. In two-generation families haplotypes may be inferred with high precision for closely linked loci even if an allele at one locus is recessive. Coding procedures are discussed and a simple solution to inferring the most likely order of three or more loci, and defining its likelihood compared with other orders, is presented.

Chromosome Mapping↗

Haplotype-based association analysis in cohort and nested case-control studies.

Genetic epidemiologic studies often collect genotype data at multiple loci within a genomic region of interest from a sample of unrelated individuals. One popular method for analyzing such data is to assess whether haplotypes, i.e., the arrangements of alleles along individual chromosomes, are associated with the disease phenotype or not. For many study subjects, however, the exact haplotype configuration on the pair of homologous chromosomes cannot be derived with certainty from the available locus-specific genotype data (phase ambiguity). In this article, we consider estimating haplotype-specific association parameters in the Cox proportional hazards model, using genotype, environmental exposure, and the disease endpoint data collected from cohort or nested case-control studies. We study alternative Expectation-Maximization algorithms for estimating haplotype frequencies from cohort and nested case-control studies. Based on a hazard function of the disease derived from the observed genotype data, we then propose a semiparametric method for joint estimation of relative-risk parameters and the cumulative baseline hazard function. The method is greatly simplified under a rare disease assumption, for which an asymptotic variance estimator is also proposed. The performance of the proposed estimators is assessed via simulation studies. An application of the proposed method is presented, using data from the Alpha-Tocopherol, Beta-Carotene Cancer Prevention Study.

Algorithms↗

Human F7 sequence is split into three deep clades that are related to FVII plasma levels.

It is widely accepted that FVII levels are strongly, consistently, and independently related to cardiovascular risk. These levels are influenced by genetic and environmental factors. Among the genetic factors, only a limited number of polymorphisms in the F7 gene have been reported, and they explain only a small proportion of the genetic variability. Recently, we have accomplished the complete dissection of the F7 quantitative trait locus responsible for all of the genetic variability observed in FVII levels. Now, we present the thorough study of the haplotype organization of F7 DNA sequence variation among individuals and the evolutionary processes that produced this variation, by sequencing 15 kb of genomic DNA sequence from the F7 locus in 40 unrelated individual (80 chromosomes) from the genetic analysis of idiopathic thrombophilia (GAIT) project as well as four non-human primate species. Our study revealed 49 polymorphisms, of which 39 SNPs were further considered. Genotyping of these DNA variations in the whole family-based GAIT sample helped resolve linkage phases, and a total of 37 distinct haplotypes were identified.Tajima's D was significantly positive in this sample, suggesting balancing selection. This parameter was a reflection of the phylogenetic structure of F7 haplotype, which was deeply split into three well-supported clades or haplogroups, suggesting that functional differences among F7 variants do not depend on a few single-site variations. Moreover, haplogroup 2 was associated with high FVII levels and haplogroup 3 with low levels. In this study, we have for the first time established a clear relation between genotypic variability structure and phenotypic variability of a particular quantitative trait involved in a complex disease.

Animals↗

Haplotype structure and population genetic inferences from nucleotide-sequence variation in human lipoprotein lipase.

Allelic variation in 9.7 kb of genomic DNA sequence from the human lipoprotein lipase gene (LPL) was scored in 71 healthy individuals (142 chromosomes) from three populations: African Americans (24) from Jackson, MS; Finns (24) from North Karelia, Finland; and non-Hispanic Whites (23) from Rochester, MN. The sequences had a total of 88 variable sites, with a nucleotide diversity (site-specific heterozygosity) of .002+/-.001 across this 9.7-kb region. The frequency spectrum of nucleotide variation exhibited a slight excess of heterozygosity, but, in general, the data fit expectations of the infinite-sites model of mutation and genetic drift. Allele-specific PCR helped resolve linkage phases, and a total of 88 distinct haplotypes were identified. For 1,410 (64%) of the 2,211 site pairs, all four possible gametes were present in these haplotypes, reflecting a rich history of past recombination. Despite the strong evidence for recombination, extensive linkage disequilibrium was observed. The number of haplotypes generally is much greater than the number expected under the infinite-sites model, but there was sufficient multisite linkage disequilibrium to reveal two major clades, which appear to be very old. Variation in this region of LPL may depart from the variation expected under a simple, neutral model, owing to complex historical patterns of population founding, drift, selection, and recombination. These data suggest that the design and interpretation of disease-association studies may not be as straightforward as often is assumed.

Animals↗

Simple estimates of haplotype relative risks in case-control data.

Methods of varying complexity have been proposed to efficiently estimate haplotype relative risks in case-control data. Our goal was to compare methods that estimate associations between disease conditions and common haplotypes in large case-control studies such that haplotype imputation is done once as a simple data-processing step. We performed a simulation study based on haplotype frequencies for two renin-angiotensin system genes. The iterative and noniterative methods we compared involved fitting a weighted logistic regression, but differed in how the probability weights were specified. We also quantified the amount of ambiguity in the simulated genes. For one gene, there was essentially no uncertainty in the imputed diplotypes and every method performed well. For the other, approximately 60% of individuals had an unambiguous diplotype, and approximately 90% had a highest posterior probability greater than 0.75. For this gene, all methods performed well under no genetic effects, moderate effects, and strong effects tagged by a single nucleotide polymorphism (SNP). Noniterative methods produced biased estimates under strong effects not tagged by an SNP. For the most likely diplotype, median bias of the log-relative risks ranged between -0.49 and 0.22 over all haplotypes. For all possible diplotypes, median bias ranged between -0.73 and 0.08. Results were similar under interaction with a binary covariate. Noniterative weighted logistic regression provides valid tests for genetic associations and reliable estimates of modest effects of common haplotypes, and can be implemented in standard software. The potential for phase ambiguity does not necessarily imply uncertainty in imputed diplotypes, especially in large studies of common haplotypes.

Algorithms↗

Investigations of HLA-F and HLA-G 3'UTR Polymorphisms in Preeclampsia and Fetal Growth Restriction Indicate a Possible Role of HLA-F-HLA-G Haplotypes and Diplotypes.

HLA-F and HLA-G may be involved in the pathogeneses of preeclampsia and fetal growth restriction (FGR). However, the functions of HLA-F and HLA-G in placental dysfunction remain unclear. The aim was to investigate differences in the prevalence of specific HLA-F and HLA-G gene allelic polymorphisms, genotypes, haplotypes, and diplotypes between controls and cases with preeclampsia or FGR. In total, blood samples from 365 pregnant females (controls, n&#x2009;=&#x2009;192; preeclampsia, n&#x2009;=&#x2009;164; FGR, n&#x2009;=&#x2009;19) in their second and third trimester, and corresponding cordial blood samples (reflecting newborns, n&#x2009;=&#x2009;160) were obtained after delivery. Genomic DNA was sequenced with a focus on the specific gene polymorphisms in the HLA-F gene locus, especially the single nucleotide polymorphisms (SNPs) rs1362126 (G/A), rs2523405 (T/G) and rs2523393 (A/G), as well as the rs371194629 (14-bp ins/del) in the 3'UTR of HLA-G. Haplotype and diplotype distributions were obtained using PHASE v2.1, and linkage disequilibrium analyses were performed. SNPs in the HLA-F gene locus and the 3'UTR of HLA-G were not associated with the risk of preeclampsia or FGR. The SNPs did not correlate with fetal-placental weight ratio, deviation of birth weight at gestational age, and placental weight. However, a trend towards an absence of certain HLA-F-HLA-G extended diplotypes in preeclampsia was observed. The current study does not support associations of the investigated HLA-F SNPs with preeclampsia or FGR. However, further studies are needed to evaluate the possible role of certain fetal HLA-F-HLA-G extended haplotypes and diplotypes in preeclampsia.

Humans↗

Recombination rates across the HLA complex: use of microsatellites as a rapid screen for recombinant chromosomes.

Meiotic recombination does not appear to occur randomly across chromosomes, but rather seems to be restricted to specific regions. A striking example of this phenomenon is illustrated by the HLA class II region. No recombination within the 100 kb encompassing the DRB1-DQA1-DQB1 loci has been reported, whereas the random association of TAP1 with TAP2 alleles suggests the presence of a hotspot for recombination within the 15 kb separating the closest variant sites of these two loci. Recombination rates between loci may provide clues to the functional properties of haplotypes. Absence of recombination may suggest the necessity to keep alleles of certain genes in phase and, alternatively, high recombination rates may suggest selective pressure to diversify haplotypes within the population. To address this issue, recombination rates across the HLA complex were determined using the 59 Centre d'Etude Polymorphisme Humain (CEPH) pedigrees. The allele frequencies of four microsatellite markers which map at sites ranging from the telomeric to centromeric ends of the complex were determined and the markers were used as a rapid means for identification of recombinant chromosomes. Typing these as well as other polymorphic loci within the HLA class I, II and III regions allowed assignment of the segments where recombination occurred. Recombination rates within the class II region (defined here as DRB1 to DPB1) and class III region (defined here as HLA-B to DRB1) regions were 0.74% and 0.94%, respectively, both of which are within an expected range given the standard of 1% recombination rate per megabase of DNA per meiosis.(ABSTRACT TRUNCATED AT 250 WORDS)

Alleles↗

Sampling among haplotype resolutions in a coalescent-based genealogy sampler.

Analysis of the coalescent structure of a population may provide information useful in mapping disease loci. Current coalescent-based genealogy samplers require haplotyped data, but haplotypes are not always available, and it is not practical to sum over all haplotype assignments for large data sets. We describe a method of adding haplotype re-evaluation to the sampler, so that it samples not only among genealogies explaining a given haplotype configuration, but also among different haplotype configurations. Several different haplotype-rearrangement strategies are considered, but the simplest-inverting the phase of a single site in a single individual-appears to be the most successful. The straightforward haplotype sampler does not mix well; heating approaches can greatly improve its performance.

Computer Simulation↗

Arthritis in mice: allogeneic pregnancy protects more than syngeneic by attenuating cellular immune response.

OBJECTIVE: We tested the hypothesis that collagen induced arthritis benefits more from allogeneic pregnancy than syngeneic pregnancy. METHODS: Arthritis was induced in female B10.RIII (H-2r) mice by injecting bovine type II collagen. Female mice were subsequently paired, one group with q-haplotype males (B10.Q) and the other with r-haplotype males (B10.RIII). The effect of q- and r-haplotype was measured by determining the acute phase reactant serum amyloid A (m-SAA), bovine anti-collagen type II antibodies (a-CBII), and the ratio of CD4/CD8 T lymphocytes during pregnancy and after delivery. Clinical assessment of arthritis was also performed. RESULTS: The number of mice with maximum severity of clinical arthritis was significantly higher in the syngeneic group (11/20 vs 5/21; p = 0.04). Although we noted that in the allogeneic group the females had had a significantly higher level of a-CBII during pregnancy (p = 0.02), we also found that the ratio of CD4/CD8 was higher in the syngeneic group even if it was measured during (p = 0.04) or after gestation (p = 0.05). Taking into account all the cases of arthritis initiated in the post-gestational period there was no difference in m-SAA or in a-CBII between the 2 groups, but the ratio of CD4/CD8 was higher in the syngeneic group measured during (p = 0.03) or post gestation (p = 0.02). CONCLUSION: Allogeneic pregnancy benefits more than syngeneic pregnancy by attenuating the cellular immune response, and the ratio of CD4/CD8 indicates the attenuation of cellular immunity when measured during gestation or post partum.

Animals↗

Haplotype-based association analysis in cohort studies of unrelated individuals.

Exploring the associations between haplotypes and disease phenotypes is an important step toward the discovery of genes that influence complex human diseases. When unrelated subjects are sampled, haplotypes are often ambiguous because of the unknown gametic phase of the measured sites along a chromosome. We consider cohort studies of unrelated subjects which collect data on potentially censored ages of onset of disease along with unphased genotypes and possibly time-varying environmental factors. We formulate the effects of haplotypes and environmental variables on the time to disease occurrence through a semiparametric Cox proportional hazards model, which can accommodate a variety of genetic mechanisms as well as gene-environment interactions. We develop a simple and fast expectation-maximization algorithm to maximize the likelihood for the relative risks and other parameters based on the observable data of unphased genotypes and potentially censored ages of onset. The resultant estimators are consistent, efficient, and asymptotically normal. Simulation studies show that, for practical situations, the parameter estimators are virtually unbiased, the association tests maintain type I errors near nominal levels, the confidence intervals have proper coverage probabilities, and the efficiency loss due to unknown gametic phase is small.

Algorithms↗