Search PubMedSearch

SEARCH · Search PubMed

Results for “Transcriptome Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

Predicted brain-regional gene expression patterns in individuals living with Alzheimer's disease.

Studying brain gene expression in Alzheimer's Disease (AD) remains difficult as postmortem brain is difficult to access, cannot be used to guide donor treatment, may be confounded by environmental factors before and after death, and is difficult to link to early AD states or disease progression. To circumvent these limitations, several studies have tested blood transcriptome biomarkers for AD. However, gene-expression levels in the blood have limited correlation with those in the brain. To evaluate the potential of monitoring Alzheimer's progression with peripheral data, we used transcriptome-imputation to identify brain-region-specific AD-associated gene-expression differences in cohorts with blood-based transcriptome data. This approach provides a high-resolution image of AD-associated molecular differences in the brains of individuals actively living with disease. We analyzed eight AD studies (777 AD cases, 779 cognitively unimpaired controls), imputing transcriptomes in 10 brain regions via the Brain Gene Expression and Network Imputation Engine (BrainGENIE). Hundreds of differentially expressed genes (DEGs) associated with AD were identified in nine brain regions, with anterior cingulate cortex and amygdala showing the most differential expression. AD-associated genes were enriched in pathways such as proteostasis, mitochondrial dysfunction, and immune activation. We observed significant yet moderate concordance between imputed AD-associated changes and those directly measured in the dorsolateral prefrontal cortex and cerebellum. These transcriptomic changes can guide future in vitro studies focused on pathogenesis or be targets of novel therapeutic development. In conclusion, we demonstrated the scope and utility of brain expression imputation from the peripheral transcriptome, laying the groundwork for biomarker discovery and prospective AD studies.

Alzheimer Disease

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r = 0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

Reimagining research papers as interactive and reliable AI agents.

Here we introduce Paper2Agent, an automated framework that converts research papers into artificial intelligence (AI) agents. Paper2Agent transforms research output from passive artefacts into active systems that accelerate use and discovery. Conventional research papers require readers to understand and adapt the paper's code, data and methods to their work, creating barriers to dissemination and reuse. Paper2Agent addresses this challenge by converting a paper into an AI agent that functions as a virtual corresponding author, exposing its manuscript, supplementary materials, datasets, code and workflows as active, agent-native knowledge rather than static text. It analyses the paper and codebase using multiple agents to construct a model context protocol (MCP) server, then generates and runs tests to refine and increase robustness of the MCP. These paper MCPs can be connected to a chat agent (such as Claude Code) to carry out complex scientific queries through natural language while invoking tools and workflows from the paper. We demonstrate Paper2Agent's effectiveness through case studies. Paper2Agent created an agent that leveraged AlphaGenome1 to interpret genomic variants and agents based on Scanpy2 and TISSUE (transcript imputation with spatial single-cell uncertainty estimation)3 to conduct single-cell and spatial transcriptomics analyses. We validate that these agents reproduce the results of the original papers and carry out novel user queries. Paper2Agent created multiple agents that collaborate to prioritize a causal gene for psoriasis. By turning static papers into interactive AI agents, Paper2Agent introduces a paradigm for knowledge dissemination and a collaborative ecosystem of AI co-scientists.

Journal Article

OmicsPred as a centralised resource for genetic prediction of multi-omic traits.

Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.

Journal Article

Alterations in ether lipid metabolism in obesity revealed by systems genomics of multi-omics datasets.

Ratios between two metabolites are sensitive indicators of metabolic changes. Lipidomic profiling studies have revealed that plasma ether lipids, a class of glycero- and glycerophospho-lipids with reported health benefits, are negatively associated with obesity. Here, we utilized lipid ratios as surrogate markers of lipid metabolism to explore the processes underlying the inverse relationship between ether lipid metabolism and obesity. Plasma lipidomics data from two independent human cohorts (n = 10,339 and n = 4,492) were integrated to assess the associations between 82 lipid ratios and obesity-related markers in males and females. Results were externally validated using mouse transcriptomics data from the Hybrid Mouse Diversity Panel (n = 152-227 across 74 strains). Genome-wide association studies using imputed genotypes from a population cohort (n = 4,492) were performed to examine the genetic architecture of the ratios. Findings showed that waist circumference (WC), body mass index, and waist-hip ratio were inversely associated with total plasmalogens relative to total phospholipids in both sexes. Ratios comprising product-substrate pairs positioned either side of enzymes involved in plasmalogen synthesis and degradation showed positive and negative associations with WC, respectively. Branched-chain fatty acids negatively correlated with WC, while omega-6 polyunsaturated fatty acids exhibited differing associations depending on their position within the pathway. Mouse transcriptomics corroborated these results. Genomics data showed strong associations between ratios containing choline-plasmalogens and single-nucleotide polymorphisms in the transmembrane protein 229B (TMEM229B) gene region. This work demonstrates the utility of lipid ratios in understanding lipid metabolism. By applying the ratios to multi-omic datasets, we identified alterations in enzymatic activity and genetic variants likely affecting ether lipid synthesis in obesity that could not have been obtained from lipidomics data alone. Additionally, we characterized a potential role for TMEM229B, offering new perspectives on ether lipid metabolism and regulation.

Humans

New Genetic Loci Implicated in Cardiac Morphology and Function Using Three-Dimensional Population Phenotyping.

BACKGROUND: Cardiac remodeling occurs in the mature heart and is a cascade of adaptations in response to stress, which are primed in early life. A key question remains as to the processes that regulate the geometry and motion of the heart and how it adapts to stress. METHODS: We performed spatially resolved phenotyping using machine learning-based analysis of cardiac magnetic resonance imaging in 47 549 UK Biobank participants. We analyzed 16 left ventricular spatial phenotypes, including regional myocardial wall thickness and systolic strain in both circumferential and radial directions. In up to 40 058 participants, genetic associations across the allele frequency spectrum were assessed using genome-wide association studies with imputed genotype participants, and exome-wide association studies and gene-based burden tests using whole-exome sequencing data. We integrated transcriptomic data from the GTEx project and used pathway enrichment analyses to further interpret the biological relevance of identified loci. To investigate causal relationships, we conducted Mendelian randomization analyses to evaluate the effects of blood pressure on regional cardiac traits and the effects of these traits on cardiomyopathy risk. RESULTS: We found 42 loci associated with cardiac structure and contractility, many of which reveal patterns of spatial organization in the heart. Whole-exome sequencing revealed 3 additional variants not captured by the genome-wide association study, including a missense variant in CSRP3 (minor allele frequency 0.5%). The majority of newly discovered loci are found in cardiomyopathy-associated genes, suggesting that they regulate spatially distinct patterns of remodeling in the left ventricle in an adult population. Our causal analysis also found regional modulation of blood pressure on cardiac wall thickness and strain. CONCLUSIONS: These findings provide a comprehensive description of the pathways that orchestrate heart development and cardiac remodeling. These data highlight the role that cardiomyopathy-associated genes have on the regulation of spatial adaptations in those without known disease.

Humans

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics

Enhancing pan-cancer spatial transcriptomics at single-cell resolution with stPainter.

Subcellular spatial transcriptomics can resolve tissue architecture at cellular scale, but sparse gene panels and limited detection sensitivity constrain downstream analysis. Existing enhancement methods often require tissue-matched single-cell RNA sequencing (scRNA-seq) references and dataset-specific retraining. Here we show that stPainter, a conditional generative model pretrained on a pan-cancer scRNA-seq atlas, can enhance spatial transcriptomics data without matched references or retraining. Using a latent diffusion architecture guided by Stochastic Differential Equations (SDE), stPainter reconstructs expanded expression profiles from sparse measurements and produces latent representations for clustering and cell-state analysis. When we apply stPainter upon 6 spatial transcriptomics datasets of different cancer types, we demonstrate that our model empowers downstream biological analyses, including fine-grained subpopulation clustering and pathway enrichment. Comparison with spatially resolved proteomics (CODEX) provided independent support for regional agreement between imputed cellular compositions and protein-level tissue organization. These results establish stPainter as a scalable approach for analyzing tumor microenvironments without auxiliary sequencing data.

Spatial Transcriptomics

Integrative cross-tissue transcriptome-wide association and metabolomic analysis reveals novel genetic risk loci for aortic aneurysm.

BACKGROUND: Aortic aneurysm (AA) is a life-threatening cardiovascular condition with a strong genetic component, however, its molecular mechanisms remain poorly understood. Although genome-wide association studies (GWAS) have identified numerous risk loci, most prior studies have investigated genetic and metabolic factors separately, leaving the causal pathways from genetic variants to disease largely unexplored. METHODS: We established an integrative framework combining cross-tissue transcriptome-wide association studies (TWAS) with metabolomic mediation analysis. First, we integrated GWAS data from FinnGen R12 with multi-tissue expression quantitative trait loci (eQTL) data from Genotype-Tissue Expression Project (GTEx) V8, then performed cross-tissue TWAS using the Unified Test for MOlecular SignaTures (UTMOST) and single-tissue validation with the Functional Summary-based Imputation (FUSION) to prioritize susceptibility genes. Second, we applied Mendelian randomization (MR), colocalization, and Fine-mapping Of CaUsal gene Sets (FOCUS) to assess causality and identify high-confidence genes. Third, we performed metabolite mediation analysis to uncover metabolic pathways linking genetic variants to disease risk. Finally, we validated key findings in mouse models of thoracic aortic aneurysm (TAA) and abdominal aortic aneurysm (AAA) using Quantitative Real-Time Reverse Transcription Polymerase Chain Reaction (RT-qPCR) and Western blotting. RESULTS: We identified multiple novel susceptibility genes for AA and its subtypes. Key genes included ADH family members (ADH1A, ADH1B, ADH4, ADH6) and ZNF827, which showed cross-subtype associations with strong colocalization evidence in vascular tissues. Metabolite mediation analysis revealed significant pathways involving N-acetylphenylalanine and methionine sulfoxide. Functional enrichment revealed distinct biological mechanisms: AA and AAA were primarily associated with metabolic pathways, whereas TAA-related genes were enriched in developmental and contractile processes. PheWAS indicated no significant off-target associations. Critically, experimental validation in mouse models confirmed significant upregulation of ZNF827 in TAA and ADH6 in AAA at both mRNA and protein levels, corroborating the genetic predictions. CONCLUSION: This integrated cross-omics analysis identifies novel genetic loci and, crucially, uncovers specific nutrient-related metabolic pathways that mediate genetic risk. These findings provide a mechanistic basis for future nutritional and metabolic intervention studies in AA and its subtypes.

MAGMA

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics