Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genotype imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Variation of the N-acetyltransferase 2 gene in a Romanian and a Kyrgyz population.

As part of a project on environmental disasters in minority populations, this study aimed to evaluate differences in the sequence of N-acetyltransferase 2 (NAT2) as a metabolic susceptibility gene in yet unexplored ethnicities. Eight single nucleotide polymorphisms (SNP) in the NAT2 coding region and a variant in the 3' flanking region were analyzed in 290 unrelated Kyrgyz and 140 unrelated Romanians by SNP-specific PCR analysis. The variants 341C, 481T, and 803G were less and 857A more prevalent in Kyrgyz (P < 0.0001). The variant at site 857 indicates Asian descent. 282C>T and 590G>A showed no significant variation by ethnicity. 364G>A and 411A>T turned out to be monomorphic. Database comparisons of the NAT2 minor allele frequencies support that Romanians belong to Caucasians and Kyrgyz are in between Caucasians and East Asians. The distributions of predicted haplotypes differed significantly between the two ethnicities where the Kyrgyz showed a higher genetic diversity. The haplotype without mutations was more common in Kyrgyz (40.1% in Kyrgyz, 29.3% in Romanians). Accordingly, the imputed slow acetylator phenotype was less prevalent in Kyrgyz (35.2% versus 51.4% in Romanians). We found pronounced ethnic differences in NAT2 genotypes with yet unknown effect on the health risks for environmental or occupational exposures in minority populations.

Acetylation↗

Modeling and E-M estimation of haplotype-specific relative risks from genotype data for a case-control study of unrelated individuals.

The US National Cancer Institute has recently sponsored the formation of a Cohort Consortium (http://2002.cancer.gov/scpgenes.htm) to facilitate the pooling of data on very large numbers of people, concerning the effects of genes and environment on cancer incidence. One likely goal of these efforts will be generate a large population-based case-control series for which a number of candidate genes will be investigated using SNP haplotype as well as genotype analysis. The goal of this paper is to outline the issues involved in choosing a method of estimating haplotype-specific risk estimates for such data that is technically appropriate and yet attractive to epidemiologists who are already comfortable with odds ratios and logistic regression. Our interest is to develop and evaluate extensions of methods, based on haplotype imputation, that have been recently described (Schaid et al., Am J Hum Genet, 2002, and Zaykin et al., Hum Hered, 2002) as providing score tests of the null hypothesis of no effect of SNP haplotypes upon risk, which may be used for more complex tasks, such as providing confidence intervals, and tests of equivalence of haplotype-specific risks in two or more separate populations. In order to do so we (1) develop a cohort approach towards odds ratio analysis by expanding the E-M algorithm to provide maximum likelihood estimates of haplotype-specific odds ratios as well as genotype frequencies; (2) show how to correct the cohort approach, to give essentially unbiased estimates for population-based or nested case-control studies by incorporating the probability of selection as a case or control into the likelihood, based on a simplified model of case and control selection, and (3) finally, in an example data set (CYP17 and breast cancer, from the Multiethnic Cohort Study) we compare likelihood-based confidence interval estimates from the two methods with each other, and with the use of the single-imputation approach of Zaykin et al. applied under both null and alternative hypotheses. We conclude that so long as haplotypes are well predicted by SNP genotypes (we use the Rh2 criteria of Stram et al. [1]) the differences between the three methods are very small and in particular that the single imputation method may be expected to work extremely well.

Algorithms↗

Inference and visualization of complex genotype-phenotype maps with gpmap-tools.

Understanding how biological sequences give rise to observable traits, that is, how genotype maps to phenotype, is a central goal in biology. Yet our knowledge of genotype-phenotype maps in natural systems is limited due to the high dimensionality of sequence space and the context-dependent effects of mutations. The emergence of Multiplex assays of variant effect (MAVEs), along with large collections of natural sequences, offer new opportunities to empirically characterize these maps at an unprecedented scale. However, tools for statistical and exploratory analysis of these high-dimensional data are still needed. To address this gap, we developed gpmap-tools (https://github.com/cmarti/gpmap-tools), a python library that integrates a series of models for inference, phenotypic imputation, and error estimation from MAVE data or collections of natural sequences in the presence of genetic interactions of every possible order. gpmap-tools also provides methods for summarizing patterns of epistasis and visualization of genotype-phenotype maps containing up to millions of genotypes. To demonstrate its utility, we used gpmap-tools to infer genotype-phenotype maps containing 262,144 variants of the Shine-Dalgarno sequence from both genomic 5'UTR sequences and experimental MAVE data. Visualization of the inferred landscapes consistently revealed high-fitness ridges that link core motifs at different distances from the start codon. In summary, gpmap-tools provides a flexible, interpretable framework for studying complex genotype-phenotype maps, opening new avenues for understanding the architecture of genetic interactions and their evolutionary consequences.

Gaussian process↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50&#xa0;K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40&#xa0;kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Influence of leukotriene pathway polymorphisms on response to montelukast in asthma.

RATIONALE: Interpatient variability in montelukast response may be related to variation in leukotriene pathway candidate genes. OBJECTIVE: To determine associations between polymorphisms in leukotriene pathway candidate genes with outcomes in patients with asthma receiving montelukast for 6 mo who participated in a clinical trial. METHODS: Polymorphisms were typed using Sequenom matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass array spectrometry and published methods; haplotypes were imputed using single nucleotide polymorphism-expectation maximization (SNP-EM). Analysis of variance and logistic regression models were used to test for changes in outcomes by genotype. In addition, chi(2) and likelihood ratio tests were used to test for differences between groups. Case-control comparisons were analyzed using the SNP-EM Omnibus likelihood ratio test. MEASUREMENTS: Outcomes were asthma exacerbation rate and changes in FEV(1) compared with baseline. RESULTS: DNA was collected from 252 participants: 69% were white, 26% were African American. Twenty-eight SNPs in the ALOX5, LTA4H, LTC4S, MRP1, and cysLT1R genes, and an ALOX5 repeat polymorphism were successfully typed. There were racial disparities in allele frequencies in 17 SNPs and in the repeat polymorphism. Association analyses were performed in 61 whites. Associations were found between genotypes of SNPs in the ALOX5 (rs2115819) and MRP1 (rs119774) genes and changes in FEV(1) (p < 0.05), and between two SNPs in LTC4S (rs730012) and in LTA4H (rs2660845) genes for exacerbation rates. Mutant ALOX5 repeat polymorphism was associated with decreased exacerbation rates. There was strong linkage disequilibrium between ALOX5 SNPs. Associations between ALOX5 haplotypes and risk of exacerbations were found. CONCLUSIONS: Genetic variation in leukotriene pathway candidate genes contributes to variability in montelukast response.

Acetates↗

Contribution of alpha- and beta-defensins to lung function decline and infection in smokers: an association study.

BACKGROUND: Alpha-defensins, which are major constituents of neutrophil azurophilic granules, and beta-defensins, which are expressed in airway epithelial cells, could contribute to the pathogenesis of chronic obstructive pulmonary disease by amplifying cigarette smoke-induced and infection-induced inflammatory reactions leading to lung injury. In Japanese and Chinese populations, two different beta-defensin-1 polymorphisms have been associated with chronic obstructive pulmonary disease phenotypes. We conducted population-based association studies to test whether alpha-defensin and beta-defensin polymorphisms influenced smokers' susceptibility to lung function decline and susceptibility to lower respiratory infection in two groups of white participants in the Lung Health Study (275 = fast decline in lung function and 304 = no decline in lung function). METHODS: Subjects were genotyped for the alpha-defensin-1/alpha-defensin-3 copy number polymorphism and four beta-defensin-1 polymorphisms (G-20A, C-44G, G-52A and Val38Ile). RESULTS: There were no associations between individual polymorphisms or imputed haplotypes and rate of decline in lung function or susceptibility to infection. CONCLUSION: These findings suggest that, in a white population, the defensin polymorphisms tested may not be of importance in determining who develops abnormally rapid lung function decline or is susceptible to developing lower respiratory infections.

Adult↗

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r&#x2009;=&#x2009;0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans↗

Integrative cross-tissue transcriptome-wide association and metabolomic analysis reveals novel genetic risk loci for aortic aneurysm.

BACKGROUND: Aortic aneurysm (AA) is a life-threatening cardiovascular condition with a strong genetic component, however, its molecular mechanisms remain poorly understood. Although genome-wide association studies (GWAS) have identified numerous risk loci, most prior studies have investigated genetic and metabolic factors separately, leaving the causal pathways from genetic variants to disease largely unexplored. METHODS: We established an integrative framework combining cross-tissue transcriptome-wide association studies (TWAS) with metabolomic mediation analysis. First, we integrated GWAS data from FinnGen R12 with multi-tissue expression quantitative trait loci (eQTL) data from Genotype-Tissue Expression Project (GTEx) V8, then performed cross-tissue TWAS using the Unified Test for MOlecular SignaTures (UTMOST) and single-tissue validation with the Functional Summary-based Imputation (FUSION) to prioritize susceptibility genes. Second, we applied Mendelian randomization (MR), colocalization, and Fine-mapping Of CaUsal gene Sets (FOCUS) to assess causality and identify high-confidence genes. Third, we performed metabolite mediation analysis to uncover metabolic pathways linking genetic variants to disease risk. Finally, we validated key findings in mouse models of thoracic aortic aneurysm (TAA) and abdominal aortic aneurysm (AAA) using Quantitative Real-Time Reverse Transcription Polymerase Chain Reaction (RT-qPCR) and Western blotting. RESULTS: We identified multiple novel susceptibility genes for AA and its subtypes. Key genes included ADH family members (ADH1A, ADH1B, ADH4, ADH6) and ZNF827, which showed cross-subtype associations with strong colocalization evidence in vascular tissues. Metabolite mediation analysis revealed significant pathways involving N-acetylphenylalanine and methionine sulfoxide. Functional enrichment revealed distinct biological mechanisms: AA and AAA were primarily associated with metabolic pathways, whereas TAA-related genes were enriched in developmental and contractile processes. PheWAS indicated no significant off-target associations. Critically, experimental validation in mouse models confirmed significant upregulation of ZNF827 in TAA and ADH6 in AAA at both mRNA and protein levels, corroborating the genetic predictions. CONCLUSION: This integrated cross-omics analysis identifies novel genetic loci and, crucially, uncovers specific nutrient-related metabolic pathways that mediate genetic risk. These findings provide a mechanistic basis for future nutritional and metabolic intervention studies in AA and its subtypes.

MAGMA↗

NAT2, GSTM-1, cigarette smoking, and risk of colon cancer.

Cigarette smoking has been associated inconsistently with colon cancer. The extent to which genetic profile influences susceptibility to the inducement of colon cancer by cigarette smoking is not known. In this study, we evaluated the associations between smoking cigarettes and polymorphisms of the NAT2 and GSTM-1 genes using data obtained from an incident case-control study of 1993 cases of colon cancer and 2410 age- and sex- matched controls. Neither NAT2 nor GSTM-1 polymorphisms were significantly associated with colon cancer, except among older women, in whom the intermediate/rapid imputed phenotype was associated with increased risk of colon cancer [odds ratio (OR) = 1.4, 95% confidence interval (CI) = 1.0-1.81. Using several indicators of cigarette smoking, we observed no significant interaction between these genotypes and cigarette smoking and colon cancer. The major variation in association with colon cancer was from the amount of cigarette exposure, with those smoking a pack or more of cigarettes per day being at an approximately 40% increased risk of colon cancer; this association did not vary by genotype. However, those who stopped smoking 5-14 years prior to diagnosis and who where intermediate/rapid acetylators were at a slightly greater risk than those who were slow acetylators (for men, OR = 1.6, 95% CI = 1.0-2.4; for women, OR = 2.5, 95% CI = 1.4-4.4). Associations were similar when proximal and distal tumors were examined and separated for age at the time of diagnosis. The lack of an association does not rule out the possibility of other genetic polymorphisms interacting with cigarette smoke to cause colon cancer, nor does it take into account individual phenotypic variability.

Adult↗

A note on a generalized single step theory for any number of hierarchical genomic matrices.

BACKGROUND: The Single Step algorithm allows combining information from genotyped and un-genotyped individuals, provided they are connected by a pedigree. However, current single step theory is limited to a single list of markers. RESULTS: We present a generalized single step (GSS) method that can accommodate any number of hierarchical molecular datasets (e.g. sequence, high and low density arrays) and pedigree, avoiding imputation. We prove that a similar efficient inversion algorithm exists. The method is recursive, starting with the highest marker density scenario. We illustrate the method with simulation and show that GSS can increase predictive accuracy compared to standard single step. R code is provided so that custom scenarios can be easily compared, either with simulated or real data. CONCLUSION: The method developed generalizes extant single step theory to any number of hierarchical molecular relationship matrices, broadening the scenarios where single step can be applied. A topic of particular interest can be ecology field data or human populations where pedigree is not available, but where samples sequenced and genotyped at different densities can exist. GSS can also be a useful tool to optimize allocation of genotyping and / or sequencing resources.

Algorithms↗

CYClones: a highly powered, fully genotyped, eight-parent yeast mapping population.

The budding yeast Saccharomyces cerevisiae is a remarkably adaptable organism that thrives in diverse environments. Global sequencing of natural isolates has revealed extensive genetic diversity within the species. Here, we describe the construction and characterization of CYClones (Collaborative Yeast Cross clones), a library of 11,392 segregants generated from a multiparent funnel cross of eight genetically diverse parental strains. To enable the genetic dissection of complex traits, we imputed whole-genome sequences for all segregants and show that CYClones captures a substantial fraction of the global genetic diversity of S. cerevisiae. Haplotype representation is well maintained, with each parental haplotype present at >5% frequency across >95% of the genome. Simulations demonstrate that CYClones has &#x2265;95% power to detect variants with heritability as low as 0.36%, with mapping resolution often finer than the length of a single gene. In summary, CYClones is a powerful community resource for dissecting the genetic architecture of complex and quantitative traits, uncovering context-dependent mutational effects, and identifying causal variants underlying phenotypic diversity.

Saccharomyces cerevisiae↗

Towards a Standard Threshold for Genome Wide Significance in Dogs.

Genome-wide association studies (GWAS) are a foundational step in tying phenotype to genotype, relying on statistical significance thresholds to distinguish true- from false-positive signals of association. Dog genomics has long relied on per-study Bonferroni thresholds of significance, basing these on SNP chip levels of markers (~100&#x2009;k to >&#x2009;14&#x2009;M variable sites). However, as the field progresses into whole genome imputation analyses and more powerful meta-analyses, there is a clear need to develop a standard significance threshold for common-variant GWAS. Using 1591 dogs from the broad-ancestry Dog10K dataset, we performed permutation analysis and developed GWAS thresholds for datasets using either 1% or 5% minor allele frequencies. The resultant p-values, 4.2&#x2009;&#xd7;&#x2009;10-7 and 5.0&#x2009;&#xd7;&#x2009;10-7 respectively, are similar to previous Bonferroni levels (p-value ~6&#x2009;&#xd7;&#x2009;10-7), but less restrictive than the standard human p-value, 5&#x2009;&#xd7;&#x2009;10-8, which is sometimes used in dog studies. Given the diverse haplotypes from the >&#x2009;320 breeds in the Dog10K input dataset, we suggest a p-value of 4&#x2009;&#xd7;&#x2009;10-7 as a standard significance threshold that could be applied to any dog GWAS.

Animals↗

Bayesian methods for quantitative trait loci mapping based on model selection: approximate analysis using the Bayesian information criterion.

We describe an approximate method for the analysis of quantitative trait loci (QTL) based on model selection from multiple regression models with trait values regressed on marker genotypes, using a modification of the easily calculated Bayesian information criterion to estimate the posterior probability of models with various subsets of markers as variables. The BIC-delta criterion, with the parameter delta increasing the penalty for additional variables in a model, is further modified to incorporate prior information, and missing values are handled by multiple imputation. Marginal probabilities for model sizes are calculated, and the posterior probability of nonzero model size is interpreted as the posterior probability of existence of a QTL linked to one or more markers. The method is demonstrated on analysis of associations between wood density and markers on two linkage groups in Pinus radiata. Selection bias, which is the bias that results from using the same data to both select the variables in a model and estimate the coefficients, is shown to be a problem for commonly used non-Bayesian methods for QTL mapping, which do not average over alternative possible models that are consistent with the data.

Alleles↗

FINEMAP-miss: fine-mapping genome-wide association studies with missing genotype information.

MOTIVATION: The most informative genome-wide association studies (GWAS) are meta-analyses that have combined multiple studies to increase the GWAS sample size. Statistical fine-mapping is a key downstream analysis of GWAS to jointly evaluate the probability of causality of all variants in a genomic region of interest. Current fine-mapping methods are miscalibrated in the meta-analysis setting due to variation in sample size across the variants. RESULTS: We introduce FINEMAP-miss, a new fine-mapping method that extends the FINEMAP model to account for variant-specific missingness. We show that FINEMAP-miss is well-calibrated in meta-analysis simulations where the standard fine-mapping fails. Compared to the summary statistics imputation approach, FINEMAP-miss provides clear improvement when the causal variants have low imputation information or when the sample size or complexity of the meta-analysis setting increase. We successfully apply FINEMAP-miss on a breast cancer GWAS meta-analysis where neither the standard fine-mapping nor the summary statistics imputation are applicable. AVAILABILITY: An open source implementation of FINEMAP-miss as an R package ("finemapmiss") is available at https://github.com/JoonasKartau/finemapmiss. The archived version of FINEMAP-miss used for this publication can be found on Zenodo at https://doi.org/10.5281/zenodo.17492622. SUPPLEMENTARY INFORMATION: is available at the journal's web site.

Genome-Wide Association Study↗

Meat consumption patterns and preparation, genetic variants of metabolic enzymes, and their association with rectal cancer in men and women.

Meat consumption, particularly of red and processed meat, is one of the most thoroughly studied dietary factors in relation to colon cancer. However, it is not clear whether meat, red meat, heterocyclic amines (HCA), or polycyclic aromatic hydrocarbons (PAH) are associated with the risk for rectal cancer. Rectal cancer cases (n = 952) and controls (n = 1205) from Utah and Northern California were recruited from a population-based case-control study between September 1997 and February 2002. Detailed in-person interviews regarding lifestyle, medical history, and diet were conducted. DNA was extracted from peripheral lymphocytes obtained from whole-blood samples, and glutathione S-transferase (GST)M1 enzyme and N-acetyl transferase (NAT)2 enzyme genotypes were assessed. Although energy and cholesterol intakes were higher among cases than controls, adjustment for confounders accounted for the differences. Increased consumption of well-done red meat [odds ratio (OR) 1.33 95% CI 0.98, 1.79] was associated with an (P = 0.04) increase in risk for rectal cancer among men. The mutagen index, calculated on the bases of reported amount, doneness, and method of cooking meat, was also positively but not significantly (P = 0.24) associated with risk of rectal cancer for men (OR 1.37 95% CI 0.98, 1.92). NAT2-imputed phenotype and GSTM1 did not consistently modify rectal cancer risk associated with meat intake. These data suggest that mutagens such as HCA that form when meat is cooked may be culpable substances in rectal cancer risk, not red meat itself.

Adult↗

Minimum-recombinant haplotyping in pedigrees.

This article presents a six-rule algorithm for the reconstruction of multiple minimum-recombinant haplotype configurations in pedigrees. The algorithm has three major features: First, it allows exhaustive search of all possible haplotype configurations under the criterion that there are minimum recombinants between markers. Second, its computational requirement is on the order of O(J(2)L(3)) in current implementation, where J is the family size and L is the number of marker loci under analysis. Third, it applies to various pedigree structures, with and without consanguinity relationship, and allows missing alleles to be imputed, during the haplotyping process, from their identical-by-descent copies. Haplotyping examples are provided using both published and simulated data sets.

Algorithms↗

Simple estimates of haplotype relative risks in case-control data.

Methods of varying complexity have been proposed to efficiently estimate haplotype relative risks in case-control data. Our goal was to compare methods that estimate associations between disease conditions and common haplotypes in large case-control studies such that haplotype imputation is done once as a simple data-processing step. We performed a simulation study based on haplotype frequencies for two renin-angiotensin system genes. The iterative and noniterative methods we compared involved fitting a weighted logistic regression, but differed in how the probability weights were specified. We also quantified the amount of ambiguity in the simulated genes. For one gene, there was essentially no uncertainty in the imputed diplotypes and every method performed well. For the other, approximately 60% of individuals had an unambiguous diplotype, and approximately 90% had a highest posterior probability greater than 0.75. For this gene, all methods performed well under no genetic effects, moderate effects, and strong effects tagged by a single nucleotide polymorphism (SNP). Noniterative methods produced biased estimates under strong effects not tagged by an SNP. For the most likely diplotype, median bias of the log-relative risks ranged between -0.49 and 0.22 over all haplotypes. For all possible diplotypes, median bias ranged between -0.73 and 0.08. Results were similar under interaction with a binary covariate. Noniterative weighted logistic regression provides valid tests for genetic associations and reliable estimates of modest effects of common haplotypes, and can be implemented in standard software. The potential for phase ambiguity does not necessarily imply uncertainty in imputed diplotypes, especially in large studies of common haplotypes.

Algorithms↗

Comparison of missing data approaches in linkage analysis.

BACKGROUND: Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. RESULTS: We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. CONCLUSION: Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.

Algorithms↗