Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

The value of relatives with phenotypes but missing genotypes in association studies for quantitative traits.

The additional statistical power of association studies for quantitative traits was derived when ungenotyped relatives with phenotypes are included in the analysis. It was shown that the extra power is a simple function of the coefficient of additive genetic relationship and the phenotypic correlation coefficient between the genotyped and ungenotyped relatives. For close relatives, such as pairs of fullsibs and identical twin pairs, gains in power in the range of 10 to 30% are achieved if only one of the pair is genotyped. The theoretical results were verified by simulations. It was shown that ignoring the error in estimating the genotype of the ungenotyped relative has little impact on the estimates and on statistical power, consistent with results from quantitative trait loci (QTL) linkage studies. For genome-wide association studies in which not all relatives with phenotypes can be genotyped, our study provides a prediction of the additional power of an analysis that includes phenotypes on ungenotyped individuals, and can be used in experimental design. We show that a two-step procedure, in which missing genotypes are imputed and subsequently an association analysis is performed, is efficient and powerful.

Genetic Predisposition to Disease↗

Genetic Architecture of Idiopathic Inflammatory Myopathies From Meta-Analyses.

OBJECTIVE: Idiopathic inflammatory myopathies (IIMs, myositis) are rare systemic autoimmune disorders that lead to muscle inflammation, weakness, and extramuscular manifestations, with a strong genetic component influencing disease development and progression. Previous genome-wide association studies identified loci associated with IIMs. In this study, we imputed data from two prior genome-wide myositis studies and analyzed the largest myositis data set to date to identify novel risk loci and susceptibility genes associated with IIMs and its clinical subtypes. METHODS: We performed association analyses on 14,903 individuals (3,206 patients and 11,697 controls) with genotypes and imputed data from the Trans-Omics for Precision Medicine reference panel. Fine-mapping and expression quantitative trait locus colocalization analyses in myositis-relevant tissues indicated potential causal variants. Functional annotation and network analyses using the random walk with restart (RWR) algorithm explored underlying genetic networks and drug repurposing opportunities. RESULTS: Our analyses identified novel risk loci and susceptibility genes, such as FCRLA, NFKB1, IRF4, DCAKD, and ATXN2 in overall IIMs; NEMP2 in polymyositis; ACBC11 in dermatomyositis; and PSD3 in myositis with anti-histidyl-transfer RNA synthetase autoantibodies (anti-Jo-1). We also characterized effects of HLA region variants and the role of C4. Colocalization analyses suggested putative causal variants in DCAKD in skin and muscle, HCP5 in lung, and IRF4 in Epstein-Barr virus (EBV)-transformed lymphocytes, lung, and whole blood. RWR further prioritized additional candidate genes, including APP, CD74, CIITA, NR1H4, and TXNIP, for future investigation. CONCLUSION: Our study uncovers novel genetic regions contributing to IIMs, advancing our understanding of myositis pathogenesis and offering new insights for future research.

Humans↗

Quality of NAT2 genotyping with restriction fragment length polymorphism using DNA isolated from frozen urine.

In large studies and under field conditions common to epidemiological research, factors outside of and inside the laboratory can introduce misclassification of genetic susceptibility markers. Few reports have been made on the accuracy of genotyping individuals using DNA extracted from frozen urine that was stored for approximately 20 years. This study was performed to determine the reproducibility and accuracy of N-acetyltransferase 2 (NAT2) genotyping by RFLP analysis using DNA from stored urine. To obtain long-term frozen urine and blood samples from the same person, the databases of two large prospective studies were linked by name and date of birth. Six polymorphisms within the coding region of NAT2 were determined in 65 urine and blood samples after which, genotypes and imputed phenotypes (rapid, slow) were derived. To test reproducibility, all of the six polymorphisms were determined twice in 47 urine-blood pairs. Reproducibility of imputed phenotypes was 91.5% in urine samples and 97.9% in blood samples. To test accuracy, results for all six polymorphisms were compared between urine and blood DNA. All of the kappa's were at least 0.85 except one. Identical results for all six polymorphisms were seen in 78.5% of urine-blood pairs. Taking blood samples as a reference standard, rapid acetylators were classified as rapid in 97% of subjects (95% confidence interval, 90-100%), and slow acetylators were classified as slow also in 97% of subjects (95% confidence interval, 91-100%), when using urine. This study shows that stored urine samples can be used for DNA genotyping in large cohort studies, when blood samples are not available.

Aged↗

The role of melanocortin-1 receptor polymorphism in skin cancer risk phenotypes.

We have examined melanocortin-1 receptor (MC1R) variant allele frequencies in the general population and in a collection of adolescent dizygotic and monozygotic twins to determine statistical associations of pigmentation phenotypes with increased skin cancer risk. This included hair and skin color, freckling, mole count and sun exposed skin reflectance. Nine variants were studied and designated as either strong R (OR = 63; 95% CI 32-140) or weak r (OR = 5; 95% CI 3-11) red hair alleles. Penetrance of each MC1R variant allele was consistent with an allelic model where effects were multiplicative for red hair but additive for skin reflectance. To assess the interaction of the brown eye color gene BEY2/OCA2 on the phenotypic effects of variant MC1R alleles we imputed OCA2 genotype in the twin collection. A modifying effect of OCA2 on MC1R variant alleles was seen on constitutive skin color, freckling and mole count. In order to study the individual effects of these variants on pigmentation phenotype we have established a series of human primary melanocyte strains genotyped for the MC1R receptor. These include strains which are MC1R wild-type consensus, variant heterozygotes, and homozygotes for strong R alleles Arg151Cys and Arg160Trp. Ultrastructural analysis demonstrated that only consensus strains contained stage III and IV melanosomes in their terminal dendrites whereas Arg151Cys and Arg160Trp homozygous strains contained only immature stage I and II melanosomes. Such genetic association studies combined with the functional analysis of MC1R variant alleles in melanocytic cells should provide a link in understanding the association between pigmentary phototypes and skin cancer risk.

Alleles↗

Imputing gene-treatment interactions when the genotype distribution is unknown using case-only and putative placebo analyses--a new method for the Genetics of Hypertension Associated Treatment (GenHAT) study.

There is a sizeable literature on methods for detecting gene-environment interaction in the framework of case-control studies, particularly with reference to the assumption of independence of genotype and exposure. In the context of a clinical trial, wherein gene-drug interactions with regard to outcomes are examined, these methods may be readily applied, as gene and drug are independent by randomization. In an active-controlled trial (experimental treatment vs standard) that has collected genotype information, gene-drug interactions can be estimated. In addition, the effect of the experimental treatment vs placebo can be imputed by using data from a historical placebo-controlled trial (standard vs placebo) if either (a) genotype information is available from the historical trial or (b) assumptions are made about the prevalence of genotype and the odds ratios of genotype and disease in the historical trial using information from other studies. Motivation for these procedures is provided by the Genetics of Hypertension Associated Treatment, a large pharmacogenetics, ancillary study of a hypertension clinical trial, and examples from published hypertension trials will be used to illustrate the methods.

Controlled Clinical Trials as Topic↗

The C523A beta2 adrenergic receptor polymorphism associates with markers of asthma severity in African Americans.

Our goal was to explore associations between ss2 adrenergic receptor polymorphisms and markers of asthma severity in African American and Caucasian patients with asthma. Polymorphisms at loci -1023, -654, -47, 46, 79, 491, and 523 were genotyped and haplotypes were imputed in 143 African Americans and 336 Caucasians. C523A genotype associated with percentage of African Americans (but not of Caucasians) having an asthma exacerbation: AA, AC, and CC genotypes were 17, 29, and 40%, respectively (p = 0.018). Symptom scores, pulmonary function, and rescue inhaler use paralleled exacerbation prevalence. We conclude the 523 A allele modifies asthma severity in African Americans.

Adult↗

An epistatic genetic basis for fluctuating asymmetry of mandible size in mice.

The genetic basis of fluctuating asymmetry (FA), or nondirectional variation in the subtle differences between left and right sides of bilateral characters, continues to be of considerable theoretical interest. FA generally has been thought to arise from random noise during development and therefore to have a largely or entirely environmental origin. Whereas additive genetic variation for FA generally has been small and often insignificant, a number of investigators have hypothesized that interactions between loci, or epistasis, significantly influence FA. We tested this hypothesis by conducting a whole-genome scan to detect any epistasis in FA of centroid size in the mandibles of more than 400 mice from an F2 intercross population formed from crossing the Large (LG/J) and Small (SM/J) inbred strains. Genotypic deviations were imputed at each site 2 cM apart on all 19 autosomes, and these and centroid size asymmetry values were used in canonical correlation analyses for each of the 171 possible pairs of 19 autosomes to identify the most probable sites for epistasis. Epistasis for centroid size asymmetry was abundant, occurring far more often than was expected by chance alone (there were 30 separate instances of epistasis at the 0.001 significance level, when only two were expected by chance alone). The contributions of epistasis from 30 pairwise combinations of loci tended to suppress the additive and dominance genetic variance, but greatly increased the epistatic genetic variance for FA in centroid size given the intermediate allele frequencies of an F2 intercross population.

Animals↗

Estimation and testing of genotype and haplotype effects in case-control studies: comparison of weighted regression and multiple imputation procedures.

A popular approach for testing and estimating genotype and haplotype effects associated with a disease outcome is to conduct a population-based case/control study, in which haplotypes are not directly observed but may be inferred probabilistically from unphased genotype data. A variety of methods exist to analyse the resulting data while accounting for the uncertainty in haplotype assignment, but most focus on the issue of testing the global null hypothesis that no genotype or haplotype effects exist. A more interesting question, once a region of disease association has been identified, is to estimate the relevant genotypic or haplotypic effects and to perform tests of complex null hypotheses such as the hypothesis that some loci, but not others, are associated with disease. Here I examine the assumptions behind, and the performance of, two classes of methods for addressing this question. The first is a weighted regression approach in which posterior probabilities of haplotype assignments are used as weights in a logistic regression analysis, generating a test based on either a weighted pseudo-likelihood, or a weighted log-likelihood. The second is a multiple imputation approach using either an improper procedure in which the posterior probabilities are used to generate replicate imputed data sets, or a proper data augmentation procedure. I compare these approaches to a simple expectation substitution (haplotype trend regression) approach. In simulations, all methods gave unbiased parameter estimation but the weighted pseudo-likelihood, expectation substitution and multiple imputation methods had superior confidence interval coverage. For the weighted pseudo-likelihood and expectation substitution methods it was necessary to estimate posterior haplotype assignment probabilities using the combined case/control data, whereas for the multiple imputation approaches it was necessary to estimate these probabilities in the case and control groups separately. Overall, multiple imputation was easiest approach to implement in standard statistical software and to extend to more complex models such as those that include gene-gene or gene-environment interactions.

Case-Control Studies↗

Associations between smoking, passive smoking, GSTM-1, NAT2, and rectal cancer.

Cigarette smoking has been identified as a risk factor for colon cancer, however, much less is known about the association between cigarette smoking and rectal cancer. The purpose of this article is to evaluate the associations between rectal cancer and active and passive cigarette smoking and other forms of tobacco use. We also evaluate how genetic variants of GSTM-1 and NAT2 alter these associations. A population-based case-control study of 952 incident rectal cancer cases and 1205 controls was conducted. Detailed tobacco use information was collected as part of an interviewer-administered questionnaire. DNA was extracted from blood to examine genetic variants of GSTM-1 and NAT2. Cigarette smoking was associated with an increased risk of rectal cancer in men [odds ratio (OR)=1.5, 95% confidence interval (CI), 1.1-2.1 for current smokers; OR=1.7, 95% CI, 1.3-2.3 for smoking >20 pack-years of cigarettes relative to never-smokers]. After adjusting for active smoking, exposure to cigarette smoke of others also was associated with increased risk among men (OR=1.5, 95% CI, 1.1-2.0). Neither GSTM-1 genotype nor NAT2-imputed phenotype was independently associated with rectal cancer. However, the risk associated with smoking cigarettes among those who were GSTM-1 null relative to those who never smoked and had the GSTM-1 present genotype was OR=2.0 (95% CI, 1.2-3.3). This interaction was of borderline significance (P=0.08). Men who had the combined GSTM-1 present genotype and who were rapid acetylators had no increased risk from cigarette smoking. There were no significant associations between cigarette smoking and rectal cancer among women. This study shows that men who smoke cigarettes, especially those who smoke >20 pack-years, are at increased risk of rectal cancer. This association may be influenced by GSTM-1 genotype. Furthermore, exposure to cigarette smoke of others may increase risk of rectal cancer among men who do not smoke.

Adult↗

The CYP1A1 genotype may alter the association of meat consumption patterns and preparation with the risk of colorectal cancer in men and women.

We hypothesized that the risk of colorectal cancer associated with meat preparation methods producing heterocyclic amines or polycyclic aromatic hydrocarbons is modified by the CYP1A1 genotype alone or in combination with the GSTM1 genotype or the NAT2 imputed phenotype. A total of 952 rectal cancer cases and 1205 controls (between September 1997 and February 2002) and 1346 colon cancer cases and 1544 controls (between October 1991 and September 1994) from Utah and Northern California were recruited from a population-based case-control study. Detailed interviews ascertained lifestyle, medical history, and diet and we extracted DNA from whole blood. Risk of colorectal cancer decreased among men with the CYP1A1 *2 any variant genotype and the lowest intake of poultry and men and women with high use of white meat drippings. Risk increased among men with the CYP1A1 *1 (no variant) allele and high white meat mutagen index, but decreased among those with the CYP1A1 *2 genotype. Risk increased with a high white meat mutagen index among women with the CYP1A1 *2 genotype and the GSTM1 present genotype. Risk of colorectal cancer decreased with the CYP1A1 *2 genotype, the NAT2 slow phenotype, and the use of white meat or its drippings. The association of risk for colorectal cancer and selected red and white meat mutagen indices and the use of white meat drippings, or fried white meat variables was more evident within select combinations of the CYP1A1 genotype and either the GSTM1 genotype or NAT2 than with the CYP1A1 alone. Genetic susceptibility may modify the associations of some meat or meat preparation factors with the risk of colorectal cancer.

Animals↗

A method for identifying genes related to a quantitative trait, incorporating multiple siblings and missing parents.

When studying either qualitative or quantitative traits, tests of association in the presence of linkage are necessary for fine-mapping. In a previous report, we suggested a polytomous logistic approach to testing linkage and association between a di-allelic marker and a quantitative trait locus, using genotyped triads, consisting of an individual whose quantitative trait has been measured and his or her two parents. Here we extend that approach to incorporate marker information from entire nuclear families. By computing a weighted score function instead of a maximum likelihood test, we allow for both an unspecified correlation structure between siblings and "informative" family size. Both this approach and our original approach allow for population admixture by conditioning on parental genotypes. The proposed method allows for missing parental genotype data through a multiple imputation procedure. We use simulations based on a population with admixture to compare our method to a popular non-parametric family-based association test (FBAT), testing the null of no association in the presence of linkage.

Alleles↗

HAPLORE: a program for haplotype reconstruction in general pedigrees without recombination.

MOTIVATION: Haplotype reconstruction is an essential step in genetic linkage and association studies. Although many methods have been developed to estimate haplotype frequencies and reconstruct haplotypes for a sample of unrelated individuals, haplotype reconstruction in large pedigrees with a large number of genetic markers remains a challenging problem. METHODS: We have developed an efficient computer program, HAPLORE (HAPLOtype REconstruction), to identify all haplotype sets that are compatible with the observed genotypes in a pedigree for tightly linked genetic markers. HAPLORE consists of three steps that can serve different needs in applications. In the first step, a set of logic rules is used to reduce the number of compatible haplotypes of each individual in the pedigree as much as possible. After this step, the haplotypes of all individuals in the pedigree can be completely or partially determined. These logic rules are applicable to completely linked markers and they can be used to impute missing data and check genotyping errors. In the second step, a haplotype-elimination algorithm similar to the genotype-elimination algorithms used in linkage analysis is applied to delete incompatible haplotypes derived from the first step. All superfluous haplotypes of the pedigree members will be excluded after this step. In the third step, the expectation-maximization (EM) algorithm combined with the partition and ligation technique is used to estimate haplotype frequencies based on the inferred haplotype configurations through the first two steps. Only compatible haplotype configurations with haplotypes having frequencies greater than a threshold are retained. RESULTS: We test the effectiveness and the efficiency of HAPLORE using both simulated and real datasets. Our results show that, the rule-based algorithm is very efficient for completely genotyped pedigree. In this case, almost all of the families have one unique haplotype configuration. In the presence of missing data, the number of compatible haplotypes can be substantially reduced by HAPLORE, and the program will provide all possible haplotype configurations of a pedigree under different circumstances, if such multiple configurations exist. These inferred haplotype configurations, as well as the haplotype frequencies estimated by the EM algorithm, can be used in genetic linkage and association studies. AVAILABILITY: The program can be downloaded from http://bioinformatics.med.yale.edu.

Algorithms↗

GSTM-1 and NAT2 and genetic alterations in colon tumors.

OBJECTIVE: Phase II metabolizing enzymes such as glutathione S-transferases and N-acetyltransferase are involved in the detoxification of carcinogens. Genetic variants of genes coding for these enzymes have been evaluated as to their association with colon cancer, both as independent risk factors and as effect modifiers for associations with diet and cigarette smoking. In this study, we evaluate associations between the GSTM-1 genotype and the NAT2-imputed phenotype and acquired mutations in tumors. METHODS: Data is taken from a set of 1836 cases and 1958 controls with colon cancer who were part of a large case-control study of colon cancer and whose tumors were previously analyzed for Ki-ras, p53, and microsatellite instability (MSI). We also evaluate the modifying effects of these genetic variants with diet and cigarette smoking, factors previously identified as being associated with specific tumor alterations. RESULTS: Neither GSTM-1 nor the NAT2-imputed phenotype was independently associated with Ki-ras, p53, or MSI. Cigarette smoking significantly increased the risk of tumors involving the MSI pathway. Additionally, cigarette smoking doubled the risk of p53 transversion mutations among those who were GSTM-1 present. Cases were slightly more likely to have a p53 mutation if they frequently consumed red meat and had the imputed NAT2 intermediate/rapid phenotype relative to slow phenotype/infrequent consumers of red meat (OR 2.0, 95% CI 1.3-3.0 for intermediate/rapid). CONCLUSIONS: These data provide support that diet and cigarette smoking may be associated with specific disease pathways, although GSTM-1 and NAT2 do not independently appear to alter susceptibility to these diet and lifestyle factors.

Adult↗

Association testing with Mendel.

This report presents an overview of association testing strategies from a user's perspective, with particular attention to the capabilities of the computer program Mendel. Association testing is driven by the nature of the study sample, the nature of the disease trait, and the kind of markers employed. The practicing statistician must also choose whether to conduct parametric or nonparametric tests. Because of the complexities involved, Mendel offers users several analysis options. The different options are tied together by shared input and output conventions and a shared language for defining models. Mendel also features new statistics and theory found in no other genetics software. The most important innovations include: association testing by penetrance estimation, expansion of matched-pair designs to permutation unit designs, and a rigorous implementation of the measured genotype approach for quantitative trait loci. This report explains how Mendel imputes allele counts and conducts both asymptotic and permutation tests in the measured genotype framework.

Analysis of Variance↗

Summary report: Missing data and pedigree and genotyping errors.

Genetic epidemiology is faced with mapping complex traits to genes with relatively small effects whose phenotypes may be modulated by temporal factors. To do this, detailed and accurate data must be available on families, perhaps collected over time. The Framingham Heart Study data supplied to Genetic Analysis Workshop 13 (GAW13), along with its simulated counterpart, contain longitudinal measurements and genomic scan data on 2,885 individuals in 330 families, and offer an opportunity to examine data quality and completeness issues as they affect analytical conclusions. Six GAW13 contributions applied methods to deal with missing data, both phenotypic and genotypic, at a single time point and longitudinally, and with possible errors in pedigree structure and genotypes. The methods included missing phenotypic data imputation by Markov chain Monte Carlo sampling, propensity scoring, regression, and adjusted mean values, as well as the assessment of transmission-disequilibrium tests when missing marker data may be allele-specific. Pedigree structural errors were found by genome-wide allele-sharing probabilities, while Mendelian consistent genotype errors were evaluated through likelihoods of double-recombination events. Each of the methods reviewed here offered insights into how to better take advantage of large, time-dependent, familial data sets. However, no one of them dealt with the longitudinal and familial aspects simultaneously. Overall, more consideration needs to be given to the effects that missing data and data errors have on our ability to map complex traits efficiently and accurately.

Cardiovascular Diseases↗

Accounting for decay of linkage disequilibrium in haplotype inference and missing-data imputation.

Although many algorithms exist for estimating haplotypes from genotype data, none of them take full account of both the decay of linkage disequilibrium (LD) with distance and the order and spacing of genotyped markers. Here, we describe an algorithm that does take these factors into account, using a flexible model for the decay of LD with distance that can handle both "blocklike" and "nonblocklike" patterns of LD. We compare the accuracy of this approach with a range of other available algorithms in three ways: for reconstruction of randomly paired, molecularly determined male X chromosome haplotypes; for reconstruction of haplotypes obtained from trios in an autosomal region; and for estimation of missing genotypes in 50 autosomal genes that have been completely resequenced in 24 African Americans and 23 individuals of European descent. For the autosomal data sets, our new approach clearly outperforms the best available methods, whereas its accuracy in inferring the X chromosome haplotypes is only slightly superior. For estimation of missing genotypes, our method performed slightly better when the two subsamples were combined than when they were analyzed separately, which illustrates its robustness to population stratification. Our method is implemented in the software package PHASE (v2.1.1), available from the Stephens Lab Web site.

Algorithms↗

Polymorphisms in signal transducer and activator of transcription 3 and lung function in asthma.

BACKGROUND: Identifying genetic determinants for lung function is important in providing insight into the pathophysiology of asthma. Signal transducer and activator of transcription 3 is a transcription factor latent in the cytoplasm; the gene (STAT3) is activated by a wide range of cytokines, and may play a role in lung development and asthma pathogenesis. METHODS: We genotyped six single nucleotide polymorphisms (SNPs) in the STAT3 gene in a cohort of 401 Caucasian adult asthmatics. The associations between each SNP and forced expiratory volume in 1 second (FEV1), as a percent of predicted, at the baseline exam were tested using multiple linear regression models. Longitudinal analyses involving repeated measures of FEV1 were conducted with mixed linear models. Haplotype analyses were conducted using imputed haplotypes. We completed a second association study by genotyping the same six polymorphisms in a cohort of 652 Caucasian children with asthma. RESULTS: We found that three polymorphisms were significantly associated with baseline FEV1: homozygotes for the minor alleles of each polymorphism had lower FEV1 than homozygotes for the major alleles. Moreover, these associations persisted when we performed an analysis on repeated measures of FEV1 over 8 weeks. A haplotypic analysis based on the six polymorphisms indicated that two haplotypes were associated with baseline FEV1. Among the childhood asthmatics, one polymorphism was associated with both baseline FEV1 and the repeated measures of FEV1 over 4 years. CONCLUSION: Our results indicate that genetic variants in STAT3, independent of asthma treatment, are determinants of FEV1 in both adults and children with asthma, and suggest that STAT3 may participate in inflammatory pathways that have an impact on level of lung function.

Adult↗

Bayesian modelling of multivariate quantitative traits using seemingly unrelated regressions.

We investigate a Bayesian approach to modelling the statistical association between markers at multiple loci and multivariate quantitative traits. In particular, we describe the use of Bayesian Seemingly Unrelated Regressions (SUR) whereby genotypes at the different loci are allowed to have non-simultaneous effects on the phenotypes considered with residuals from each regression assumed correlated. We present results from simulations showing that, under rather general conditions that are likely to hold in real situations, the Bayesian SUR approach has increased probability of selecting the true model compared to univariate analyses. Finally, we apply our methods to data from subjects genotyped for 12 SNPs in the apolipoprotein E (APOE) gene. Phenotypes relate to response to treatment with atorvastatin and include changes in total cholesterol, low-density lipoprotein cholesterol, and triglycerides. Missing genotype data are naturally accommodated in our Bayesian framework by imputing them using a nested haplotype phasing algorithm.

Algorithms↗