Search PubMedSearch

SEARCH · Search PubMed

Results for “imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

MetaGLIMPSE: Meta-imputation of low-coverage sequencing data for modern and ancient genomes.

The advent of efficient and accurate imputation for low-coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between target and reference samples, leveraging imputation from multiple reference panels improves imputation accuracy; however, individual reference panel genotypes are often privacy protected. Meta-imputation bypasses individual-level data by combining single-panel imputed genotypes through estimating panel- and marker-specific weights. We present a meta-imputation method, MetaGLIMPSE, that combines estimates from multiple reference panels for low-coverage sequencing imputation. Across all our scenarios, for both modern and ancient DNA samples, MetaGLIMPSE consistently outperforms the best single-panel imputation for coverages of 0.1×-8× and across all minor-allele frequencies, equaling the combined panel imputation for some parameters. Finally, MetaGLIMPSE is computationally efficient, meta-imputing 500 whole genomes in 16% of the time of GLIMPSE2.

Humans

Pedigree-assisted genotype imputation enables cost-effective genomic prediction in Penaeus vannamei.

Genomic selection in Penaeus vannamei has long been constrained by the high cost of dense genotyping. To address this limitation, we evaluated genotype imputation from a low-density 1 K panel to a medium-density 55 K panel of the "Yellow Sea Array No. 1" and examined its impact on genomic prediction for harvest body weight in P. vannamei. A four-generation pedigree including 30 great-grandparents, 39 grandparents, 100 parents, and 608 offspring was genotyped using the 55 K panel. A two-step experimental design was implemented to (i) assess the performance of different imputation algorithms under reference population scenarios with varying proportions of siblings, and (ii) compare six alternative reference population structures incorporating parents, ancestors, and siblings. Genotype imputation using the pedigree-based method FImpute v3.0 consistently achieved higher accuracy than the population-based method Beagle v5.5. Using this pedigree-assisted approach, imputation accuracy increased from 0.73 when only parental genotypes were used to 0.84 with the inclusion of 10% siblings, and subsequently plateaued at 0.87-0.90 when sibling representation reached 20%. Across the six reference population structures, imputation accuracy was primarily driven by the availability of parental genotypes, ranging from 0.50 to 0.56 in the absence of parents to 0.88-0.89 when both parents and ancestral generations were included. Accuracy remained high when both parents were available (0.84-0.87 with siblings; 0.73 without siblings) but declined substantially when only one parent was genotyped (0.65-0.68). Imputation accuracy was positively associated with both minor allele frequency (MAF) and linkage disequilibrium (max r2LD), with LD exerting the stronger influence. Heritability estimates derived from imputed 55 K genotypes were highly consistent with those obtained from the original 55 K data (0.39 ± 0.14 vs. 0.41 ± 0.14), indicating that genotype imputation did not compromise variance component estimation. In predictive ability analyses, pedigree-based BLUP (PBLUP) achieved higher predictive ability than genomic BLUP (GBLUP) based on the 1 K panel, with predictive abilities of 0.42-0.44 for PBLUP compared with 0.34-0.35 for GBLUP. Using imputed genotypes for genomic prediction further improved predictive ability relative to the true 1 K panel, yielding values ranging from 0.35 to 0.47. Notably, when parental genotypes were included in the reference population, GBLUP based on imputed genotypes surpassed the predictive ability of PBLUP and approached that achieved with the original 55 K genotypes (0.45-0.47). Collectively, these results provide the first empirical evidence that low- to medium-density genotype imputation, combined with pedigree information, can effectively support genomic prediction in P. vannamei. This study establishes a cost-efficient and scalable framework for implementing genomic selection in P. vannamei and provides a practical reference for the application of genomic selection in other aquaculture species with constrained breeding budgets.

Animals

Effect of founder breeds on genotype imputation accuracy in Canchim cattle.

UNLABELLED: Genotype imputation is a technique used to infer unobserved genotypes based on reference panels, allowing increased marker density and cost-effective optimization for genomic selection. This study aimed to evaluate whether the inclusion of genotypes from the founder breeds Nelore (NE) and Charolais (CH) improves the imputation accuracy in the composite beef cattle breed Canchim (CA). The populations studied consisted of 804 NE, 897 CH, and 392 CA animals, all genotyped using high-density panels (777,962 SNP – single nucleotide polymorphisms). CA animals had their genotypes masked to simulate a medium-density panel (54,609 SNP). Fourteen imputation scenarios were evaluated, varying according to breed, sex, year of birth, and lineage. Imputation accuracy was determined based on the percentage of correctly imputed genotypes (PERC) and the squared Pearson’s correlation between observed and imputed genotypes (R2). PERC values ranged from 66.52% to 97.39% and R² from 0.6352 to 0.9780. The scenarios that included NE, CH, and CA (males or animals born before 2004) as the reference population for imputing CA females or CA animals born after 2004 showed the highest imputation accuracies. Therefore, the use of founder breeds in the reference population improves the accuracy of genotype imputation in CA cattle. The results indicate that a multibreed reference population, incorporating founder breeds, could provide a more robust and informative genetic basis for imputing composite cattle. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13353-026-01060-z.

Animal breeding

Adjustment for Genotype Imputation Uncertainty Corrects for Inflated Type I Error in Family-Based Association Testing.

Genotype imputation is a widely-used data augmentation approach that is applied to samples of related and/or unrelated individuals. Association testing may then be carried out on the complete data with commonly-used methods. This approach has typically not accounted for the mix of observed and imputed data, although recent work has noted the potential for introduction of confounding in case-control studies. In the Alzheimer's Disease Sequencing Project family sample we found severe inflation of the test statistics in logistic regression analysis following genotype imputation, even after standard covariate adjustments. Here we dissect sources of this inflation, which is driven by three factors: frequency-dependent bias in imputation-induced allele frequencies, differential measurement error, and differential genotyping rates in cases versus controls that introduces confounding. To address the problem, we propose a statistic, imputation deviance (), which can be easily computed from the observed and imputed genotype probabilities. We show that, as an additional fixed-effect covariate, controls the genome-wide inflation in analysis of this family-based sample, and we speculate that use of imputation deviance may also provide a practical approach to correct for genotype imputation effects in other settings, particularly when a data set is unbalanced and includes related individuals.

Humans

Multiple imputation in health-care databases: an overview and some applications.

Multiple imputation for non-response replaces each missing value by two or more plausible values. The values can be chosen to represent both uncertainty about the reasons for non-response and uncertainty about which values to impute assuming the reasons for non-response are known. This paper provides an overview of methods for creating and analysing multiply-imputed data sets, and illustrates the dramatic improvements possible when using multiple rather than single imputation. A major application of multiple imputation to public-use files from the 1970 census is discussed, and several exploratory studies related to health care that have used multiple imputation are described.

Databases, Factual

Selphi, a tool for improving genotype imputation accuracy.

Genotype imputation is a powerful tool for inferring missing genotype data in large-scale genetic studies. Over the last two decades, multiple imputation algorithms have been developed, steadily improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge, largely because existing methods rely on local haplotype matching within genomic windows and do not fully exploit the extended patterns of haplotype sharing that span entire chromosomes. Here we present Selphi, a new genotype imputation algorithm that combines the Positional Burrows-Wheeler Transform (PBWT) with a multi-stage haplotype selection heuristic operating across entire chromosomes. When compared to state-of-the-art methods Beagle 5.4, IMPUTE5, and Minimac4, Selphi showed higher accuracy on the 1000 Genomes Project and TOPMed datasets, across all super-populations and allele frequencies. Similarly, Selphi achieved higher accuracy than Beagle 5.4 on the UK Biobank dataset, which translated into improved concordance with hc-WGS GWAS summary statistics at known trait-associated loci and more accurate polygenic risk scores (PRS). Selphi outputs standard VCF files with genotype dosages (DS), haplotype-specific allele probabilities (AP1, AP2), and a per-variant dosage R-squared quality score (DR2), enabling direct integration with downstream analytical pipelines including standard post-imputation quality filtering.

Genome-Wide Association Study

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics

Evaluating the Antigen and Eplet Accuracy of DQA1 Imputations With the HaploSFHI Two-Field HLA Typing Inference Tool.

Donor/recipient mismatched HLA antigens can lead to the production of Donor-Specific Antibodies by the recipient, which are deleterious to organ transplants. The HLA-DQ locus is the most frequent target, with both the DQ beta and alpha chains involved. For deceased donors in particular, while HLA-DQB1 has been typed in emergencies for a long time, HLA-DQA1 has only recently been included. No imputation algorithmic tool was available to impute HLA-DQA1 until the development of HaploSFHI, trained on 61,393 two-field typings by NGS methods. We evaluated the accuracy of two-field HLA-DQA1 imputation from serological and two-field level HLA-A, B, DRB1, and DQB1 typings. We report a highly accurate two-field HLA-DQA1 prediction using a French test cohort of 7696 individuals, respectively reaching 92.30% and 96.45% accuracy. The average 'False Positive eplet load' stood at 0.19 and 0.07, respectively, and the average 'False Negative eplet load' at 0.18 and 0.08, respectively. A similar performance was obtained on three independent test cohorts of European ancestry (from the USA, the UK, and Portugal). Interestingly, performance was only slightly inferior on five independent test cohorts of other ethnicities (from Hong Kong and the USA) whereas it was significantly lower for two-field DRB1 imputation from its serological level. These results suggest that DQA1 can reliably be imputed even when information is totally missing, with low error risk at both antigen and eplet levels, even if the reference population is not matched. Similar additional initiatives would be welcome to confirm these findings.

Humans

The Soifua Manuia reference panel with 2,570 Samoan haplotypes improves genotype imputation quality among Samoans.

Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we used whole-genome sequencing from 1,285 Samoan individuals combined with 1000 Genomes Project (1KGP) individuals to construct an imputation reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. Here we show that this panel yielded up to two times more well-imputed (r2 ≥ 0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum; accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in accuracy over TOPMed-R3 was largest for small haplotypes, reflecting the Samoan panel's ability to capture variation not well tagged by other panels.

Haplotypes

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r = 0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans

Estimating the distribution of times from HIV seroconversion to AIDS using multiple imputation. Multicentre AIDS Cohort Study.

Multiple imputation is a model based technique for handling missing data problems. In this application we use the technique to estimate the distribution of times from HIV seroconversion to AIDS diagnosis with data from a cohort study of 4954 homosexual men with 4 years of follow-up. In this example the missing data are the dates of diagnosis with AIDS. The imputation procedure is performed in two stages. In the first stage, we estimate the residual AIDS-free time distribution as a function of covariates measured on the study participants with data provided by the participants who were seropositive at study entry. Specifically, we assume the residual AIDS-free times follow a log-normal regression model that depends on the covariates measured at enrolment on the seropositive participants. In the second stage we impute the date of AIDS diagnosis for the participants who seroconverted during the course of the study and are AIDS-free with use of the log-normal distribution estimated in the first stage and the covariates from each seroconverter's latest visit. The estimated proportions developing AIDS within 4 and within 7 years of seroconversion are 15 and 36 per cent respectively, with associated 95 per cent confidence intervals of (10, 21) and (26, 47) per cent. We discuss the Bayesian foundations of the multiple imputation technique and the statistical and scientific assumptions.

AIDS Serodiagnosis

Effects of mid-point imputation on the analysis of doubly censored data.

Doubly censored data arise in some cohort studies of the AIDS incubation period because the time of infection may be known only up to an interval defined by two successive screening tests for HIV antibody. A simple analytic approach is to impute the infection time by the mid-point of the interval and then apply standard survival techniques for right censored data. The objective of this paper is to investigate the statistical properties of such a mid-point imputation approach. We investigated the asymptotic bias of the Kaplan-Meier estimate, coverage probabilities of associated confidence intervals, bias in hazard ratio, and the size of the logrank test. We show that the statistical properties of mid-point imputation depend strongly on the underlying distributions of infection times and the incubation periods, and the width of the interval between screening tests. In the absence of treatment, the median incubation period of HIV infection is approximately 10 years, and we conclude that, for this situation, mid-point imputation is a reasonable procedure for interval widths of 2 years or less.

Bias

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

Archaic ancestry inference in imputed ancient human genomes.

When modern humans expanded from Africa into Eurasia, they interbred with archaic hominins such as Neanderthals and Denisovans. This introgression shaped human evolution, yet most insights have been gained from present-day genomes, leaving little known about how archaic variants evolved after interbreeding. Ancient genomes offer a direct view of this process, but low coverage and poor quality have limited their use. Recent advances in genotype imputation offer a way to overcome these challenges by reconstructing missing information from reference panels and recovering evolutionary signals from low-coverage data. Here, we show that imputation enables accurate detection and quantification of archaic introgression in ancient genomes, improves local archaic ancestry inference, and that regions of archaic ancestry are imputed with especially high accuracy. We further demonstrate that imputed genomes can reconstruct the trajectories of introgressed haplotypes, distinguish populations across time and geography, and identify both known and additional candidates for adaptive introgression.

Humans

slideimp: efficient imputation of DNA methylation data.

SUMMARY: We developed slideimp, an R package that extends and optimizes K-nearest neighbor (K-NN) and Principal Component Analysis (PCA) imputation with grouped and sliding-window modes for accurate and efficient imputation of microarray and whole-genome DNA methylation (DNAm) data, respectively. Under a realistic scenario, slideimp achieved ≈12-28× faster runtime and ≈3-6× peak memory usage reduction for DNAm microarray imputation (GSE286313, EPICv2, N = 72) and achieved high imputation accuracy in a whole-genome DNAm dataset (N = 41). AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/hhp94/slideimp_paper. The R package slideimp is available on CRAN (DOI: 10.32614/CRAN.package.slideimp). Version 1.0.0 of slideimp, which was used in this study, is archived on Zenodo (DOI: 10.5281/zenodo.20029382).

DNA Methylation

Strategies for the analysis of imputed data from a sample survey. The National Medical Care Utilization and Expenditure Survey.

Missing data in sample surveys is virtually unavoidable, whether it is an entire unit that is missing or only an item for a responding unit. Compensation for unit nonresponse is usually made through the assignments of weights to responding units; for item nonresponse, the compensation often is by an imputation procedure. This paper reviews the extent of missing data in a large federal survey, the National Medical Care Utilization and Expenditure Survey, and the imputation procedures used to compensate for item missing data. The effects of imputation on several types of estimates from the survey are examined. In addition, several methods for analyzing survey data with imputed values are reviewed, and recommendations about preferred strategies are made for selected circumstances.

Data Collection

The effect of imputation procedures on first birth intervals: evidence from five African fertility surveys.

In most African societies there is little motivation to remember dates of demographic events with the level of precision required in demographic surveys. Consequently it is common that the large majority of survey respondents can provide only the calendar year of occurrence or their age at the time of the event. The World Fertility Survey Group decided to handle the problem of poor date reporting by using a computer program to impute the missing information. This article illustrates the effect of these imputation procedures on cross-national differentials in the proportion of premarital first births in Benin, Cameroon, Côte d'Ivoire, Ghana, and Nigeria. The analysis demonstrates that the exceptionally low proportion of premarital first births in Ghana is an artifact of the imputation procedures.

Adolescent

A model-based approach to the imputation of missing data: home injury incidences.

Missing or incomplete data cases are a problem in all types of statistical analyses. In disease surveillance, this problem inhibits determining the actual incidence of a disease event and monitoring the disease occurrence. Several statistical techniques have been developed to impute values for incomplete data cases. We present a model-based approach to the imputation of missing data elements as applied to determining the incidence of home injury deaths.

Accidents, Home