Search PubMedSearch

SEARCH · Search PubMed

Results for “transcriptomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Revealing cancer driver genes through integrative transcriptomic and epigenomic analyses with Moonlight.

Cancer involves dynamic changes caused by (epi)genetic alterations such as mutations or abnormal DNA methylation patterns which occur in cancer driver genes. These driver genes are divided into oncogenes and tumor suppressors depending on their function and mechanism of action. Discovering driver genes in different cancer (sub)types is important not only for increasing current understanding of carcinogenesis but also from prognostic and therapeutic perspectives. We have previously developed a framework called Moonlight which uses a systems biology multi-omics approach for prediction of driver genes. Here, we present an important development in Moonlight2 by incorporating a DNA methylation layer which provides epigenetic evidence for deregulated expression profiles of driver genes. To this end, we present a novel functionality called Gene Methylation Analysis (GMA) which investigates abnormal DNA methylation patterns to predict driver genes. This is achieved by integrating the tool EpiMix which is designed to detect such aberrant DNA methylation patterns in a cohort of patients and further couples these patterns with gene expression changes. To showcase GMA, we applied it to three cancer (sub)types (basal-like breast cancer, lung adenocarcinoma, and thyroid carcinoma) where we discovered 33, 190, and 263 epigenetically driven genes, respectively. A subset of these driver genes had prognostic effects with expression levels significantly affecting survival of the patients. Moreover, a subset of the driver genes demonstrated therapeutic potential as drug targets. This study provides a framework for exploring the driving forces behind cancer and provides novel insights into the landscape of three cancer sub(types) by integrating gene expression and methylation data.

Humans

Bacteria and phage consortia modulate cecal SCFA production and host metabolism to enhance feed efficiency in ducks.

BACKGROUND: The gut microbiota influences poultry health, nutrition, feed efficiency (FE), and overall productivity. However, the relationship between gut microbes, including bacteria and phages, and FE in ducks remains underexplored. To address this, we integrated cecal 16S amplicon, metagenome, microbiota-derived short-chain fatty acids (SCFAs) profiling, liver transcriptome, and serum metabolome data to illustrate the contribution of the gut microbiome (bacteria and viruses) to duck FE. RESULTS: We reconstructed viral genomes and prokaryotic metagenome-assembled genomes (MAGs) and annotated their genes using comprehensive databases. Prokaryotic hosts of viruses were also predicted to understand virus-host dynamics within the gut ecosystem. Our results revealed that high-FE ducks have higher concentration of propionate and butyrate in cecum compared with low-FE ducks. The metagenome sequencing revealed distinct cecal microbiota profiles between two groups, with increased relative abundance of representative SCFA producers, especially Paraprevotella sp905215575 and Bacteroides sp944322345, and enhanced SCFA-biosynthesis pathways in high-FE ducks. Virome genome assembly identified two phages encoding auxiliary metabolic genes (AMGs) involved in pyruvate metabolism, enhancing nutrient availability for host bacteria to produce SCFAs (e.g., temperate phage-encoded pyruvate phosphate dikinase) or exploiting host central metabolic pathways for viral replication (e.g., lytic phage-encoded formate C-acetyltransferase). Furthermore, these representative SCFA-producing bacteria and phage consortia were associated with serum metabolites (including L-histidine and 4-hydroxydecanedioylcarnitine) linked to duck FE. CONCLUSION: Collectively, these findings provide novel insights into the gut microbial factors regulating FE in ducks, offering potential strategies to optimize poultry nutrition and productivity. Video Abstract.

Animals

NMFProfiler: a multi-omics integration method for samples stratified in groups.

MOTIVATION: The development of high-throughput sequencing enabled the massive production of "omics" data for various applications in biology. By analyzing simultaneously paired datasets collected on the same samples, integrative statistical approaches allow researchers to get a global picture of such systems and to highlight existing relationships between various molecular types and levels. Here, we introduce NMFProfiler, an integrative supervised NMF that accounts for the stratification of samples into groups of biological interest. RESULTS: NMFProfiler was shown to successfully extract signatures characterizing groups with performances comparable to or better than state-of-the-art approaches. In particular, NMFProfiler was used in a clinical study on atopic dermatitis (AD) and to analyze a multi-omic cancer dataset. In the first case, it successfully identified signatures combining known AD protein biomarkers and novel transcriptomic biomarkers. In addition, it was also able to extract signatures significantly associated to cancer survival. AVAILABILITY AND IMPLEMENTATION: NMFProfiler is released as a Python package, NMFProfiler (v0.3.0), available on PyPI.

Humans

Association Between Ticagrelor and Glucose Homeostasis Regulation: Insights from Genetic and Transcriptomic Analyses.

Emerging evidence has demonstrated the additional therapeutic benefits of ticagrelor in acute coronary syndrome (ACS) patients with diabetes. However, the underlying mechanisms of this association remain elusive. Mendelian randomization (MR) analysis using genome-wide association study (GWAS) data on ticagrelor, plasma proteomics and type 2 diabetes was employed to identify causal mediator proteins. RNA sequencing (RNA-seq) of ticagrelor-treated HepG2 cells revealed the molecular pathways regulating glucose metabolism. Genetically proxied ticagrelor was significantly associated with a reduced risk of diabetes (OR = 0.859, 95% CI: 0.783-0.934, P = 7.98E-05), and 24.41% of this effect was mediated by upregulation of BDH2 protein. In vitro experiments confirmed the enhanced effect of ticagrelor on glucose consumption. Transcriptome analysis revealed that mitochondrial respiratory chain transfer and oxidative phosphorylation (OXPHOS) were significantly enriched, and genes related to ATP biosynthesis were significantly upregulated. These findings highlight the non-platelet function of ticagrelor in maintaining glucose homeostasis, providing insights into potential drug repurposing in the future.

Humans

Transcriptomic analysis at 48 h postmortem: a proof of concept for the identification of biomarkers to estimate time since death.

BACKGROUND: The postmortem interval (PMI) refers to the time elapsed between an individual's death and the examination of the body. Tissues undergo a sequence of anatomical changes following death, which are routinely used to estimate the PMI. METHODS: To determine if these anatomical changes are associated with identifiable genomic adaptations that could characterize the PMI more accurately, we analyzed the rat skeletal muscle transcriptome at 0 and 48&#xa0;h postmortem using Clariom&#x2122; S arrays. This study investigates whether specific transcriptomic changes correlate with PMI progression, offering a potential molecular tool to complement established anatomical methods. RESULTS: A total of 3,873 differentially expressed mRNAs were identified, of which 2,787 downregulated and 1,086 upregulated transcripts. The most significantly downregulated mRNA was Tnni1 (FC = -30.95, p&#x2009;=&#x2009;1&#x2009;&#xd7;&#x2009;10-3), while the most upregulated were mt-ATP6, mt-ATP8, and mt-CO3 (FC&#x2009;>&#x2009;7.78, p&#x2009;<&#x2009;1.36&#x2009;&#xd7;&#x2009;10-12). Gene ontology (GO) enrichment analyses revealed that mRNAs upregulated at 48&#xa0;h in the PMI were primarily associated with vascular and endothelial processes, including nitric oxide transport and angiogenesis. Conversely, downregulated mRNAs were linked to mitochondrial activity and cellular metabolism, reflecting both a transient vascular response and metabolic pathway shutdown in the rat skeletal muscle. CONCLUSION: Our results demonstrate significant transcriptomic changes at 48&#xa0;h postmortem, highlighting specific genes and biological pathways that may serve as candidate biomarkers for PMI estimation.

Animals

Protocol to decode the role of transcriptionally active microbes in SARS-CoV-2-positive patients using an RNA-seq-based approach.

The elucidation of the role of microorganisms in human infections has been hindered by difficulties using conventional culture-based techniques. Here, we present a protocol for the investigation of transcriptionally active microbes (TAMs) using an RNA sequencing (RNA-seq)-based approach. We describe the steps for RNA isolation, viral genome sequencing, RNA-seq library preparation, and metatranscriptomic and transcriptomic analysis. This protocol permits a comprehensive evaluation of TAMs' contributions to the differential severity of infectious diseases, with a particular focus on diseases such as COVID-19. For complete details on the use and execution of this protocol, please refer to Devi et&#xa0;al.1.

Humans

Identification and characterization of the HSP gene family in the Chinese giant salamander: Expression patterns under combined environmental stress.

BACKGROUND: The Chinese giant salamander (Andrias davidianus) is a critically endangered living fossil species that is highly sensitive to changes in water temperature. However, systematic studies on the heat shock protein (HSP) gene family and its response mechanisms to environmental stress in this species remain limited. This study utilized transcriptome data from captive-bred salamanders exposed to combined temperature and pathogen stress. Bioinformatics tools were employed to identify the HSP gene family of A. davidianus (AndHSP) and to analyze their evolution, structure, and function, thereby revealing their regulatory mechanisms in response to environmental stress. RESULTS: A total of 72 AndHSPs were identified and classified into five subfamilies. Phylogenetic analysis revealed that each subfamily is evolutionarily conserved and functionally related. Gene expression analysis demonstrated that pathogen infection induced the expression of AndHSPs, and elevated temperature significantly intensified this response. Nine key differentially expressed genes were identified, predominantly from the AndHSP70 subfamily, with AndHSP70-18 exhibiting rapid heat-induced expression. Tissue-specific analysis showed high expression of AndHSP60 in the spleen. A qPCR validation confirmed the reliability of the transcriptome expression results. CONCLUSIONS: This study presents the first systematic identification of the AndHSP gene family and elucidates its cooperative stress response mechanisms under combined temperature and pathogen stress. These findings provide a molecular basis for understanding the species' environmental adaptation and have important implications for its conservation and artificial breeding.

Animals

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi&#x2011;omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in&#xa0;vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans

Combined somatic mutation and transcriptome analysis reveals region-specific differences in clonal architecture in human cortex.

The human cerebral cortex is specialized into regions, but little is known about how human cellular lineages shape cortical regional variation and neuronal cell-type distribution during development. Here, we map single-cell lineages of human cortical regions and neuronal subtypes using >1,000 somatic single-nucleotide variants (sSNVs) identified from deep bulk whole-genome sequencing and analyzed over 25 regions and >72,000 single cells. In the fronto-parietal cortex, sSNVs are rarely restricted, marking neuron-generating clones that disperse into neighboring regions. In contrast, the primary visual cortex harbors 30%-70% more sSNVs than the neighboring secondary visual cortex. Clones at this border exhibit more restricted dispersion, suggesting late developmental lineage segregation. Single-nucleus sSNV and whole-transcriptome analysis reveal glutamatergic neuron clones with modest regional restrictions that share low-mosaic sSNVs with some GABAergic neurons, suggesting a recent dorsal cortical progenitor. Our analysis reveals human-specific cortical lineage patterns, regional differences in clonal patterns, and late divergence of some glutamatergic/GABAergic lineages.

Humans

REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics Research.

SUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR.

Regulon

Integrative multi-omics analyses suggest a candidate microbial metabolite-associated host gene network in ulcerative colitis.

Ulcerative colitis (UC) is associated with gut microbial dysbiosis, but the host molecular alterations potentially linked to microbially derived metabolites remain incompletely understood. We integrated Mendelian randomization (MR), microbial metabolite annotation, computational target prediction, colonic transcriptomics, network analysis, and machine learning. MiBioGen microbiome GWAS data were used as exposures and FinnGen Release 12 ULCERENTER as the outcome. Metabolites linked to MR-prioritized taxa were retrieved from GutMGene, and human targets were predicted using SwissTargetPrediction and SEA. UC-related genes were defined by integrating differential expression analysis and WGCNA and then intersected with predicted metabolite targets. MR prioritized one family and eight genera showing nominal genetically supported associations with UC, but none remained significant after Benjamini-Hochberg FDR correction. Three prioritized genera were linked to 15 microbe-metabolite records, corresponding to 13 unique metabolites; nine were retained for target prediction, yielding 277 unique predicted human targets. Transcriptomic analysis identified 1,530 DEGs and a 312-gene MEgrey60 module, with 273 overlapping genes, producing 1,569 unique UC-related genes. Their intersection with the 277 predicted targets yielded 47 candidate genes. Enrichment analyses highlighted mainly metabolic and lipid-related processes. Random Forest showed the highest mean AUC across the two independent external benchmarking cohorts, and SHAP prioritized EPHX1, HSD17B2, IGFBP5, and MMP10. IBDome analysis showed inflammation-associated expression differences in these genes. This study provides a genomics-informed, hypothesis-generating framework that prioritizes candidate microbe-metabolite-host relationships in UC for future experimental validation.

Humans

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256&#x2009;Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

Spatial transcriptomics-aided localization for single-cell transcriptomics with STALocator.

Single-cell RNA-sequencing (scRNA-seq) techniques can measure gene expression at single-cell resolution but lack spatial information. Spatial transcriptomics (ST) techniques simultaneously provide gene expression data and spatial information. However, the data quality of the spatial resolution or gene coverage is still much lower than the quality of the single-cell transcriptomics data. To this end, we develop a ST-Aided Locator for single-cell transcriptomics (STALocator) to localize single cells to corresponding ST data. Applications on simulated data showed that STALocator performed better than other localization methods. When applied to the human brain and squamous cell carcinoma data, STALocator could robustly reconstruct the relative spatial organization of critical cell populations. Moreover, STALocator could enhance gene expression patterns for Slide-seqV2 data and predict genome-wide gene expression data for fluorescence in situ hybridization (FISH) and Xenium data, leading to the identification of more spatially variable genes and more biologically relevant Gene Ontology (GO) terms compared with the raw data. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

Genome-wide DNA methylation and transcriptome sequencing analyses of lens tissue in an age-related mouse cataract model.

DNA methylation is known to be associated with cataracts. In this study, we used a mouse model and performed DNA methylation and transcriptome sequencing analyses to find epigenetic indicators for age-related cataracts (ARC). Anterior lens capsule membrane tissues from young and aged mice were analyzed by MethylRAD-seq to detect the genome-wide methylation of extracted DNA. The young and aged mice had 76,524 and 15,608 differentially methylated CCGG and CCWGG sites, respectively. The Pearson correlation analysis detected 109 and 33 differentially expressed genes (DEGs) with negative methylation at CCGG and CCWGG sites, respectively, in their promoter regions. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analyses showed that DEGs with abnormal methylation at CCGG sites were primarily associated with protein kinase C signaling (Akap12, Capzb), protein threonine kinase activity (Dmpk, Mapkapk3), and calcium signaling pathway (Slc25a4, Cacna1f), whereas DEGs with abnormal methylation at CCWGG sites were associated with ribosomal protein S6 kinase activity (Rps6ka3). These genes were validated by pyrosequencing methylation analysis. The results showed that the ARC group (aged mice) had lower Dmpk and Slc25a4 methylation levels and a higher Rps6ka3 methylation than the control group (young mice), which is consistent with the results of the joint analysis of differentially methylated and differentially expressed genes. In conclusion, we confirmed the genome-wide DNA methylation pattern and gene expression profile of ARC based on the mouse cataract model with aged mice. The identified methylation molecular markers have great potential for application in the future diagnosis and treatment of ARC.

Animals

An anti-androgen resistance-related gene signature acts as a prognostic marker and increases enzalutamide efficacy via PLK1 inhibition in prostate cancer.

BACKGROUND: Anti-androgen resistance remains a major clinical challenge in the treatment of prostate cancer (PCa), leading to disease progression and treatment failure. Despite extensive research on resistance mechanisms, a reliable prognostic model for predicting patient outcomes and guiding therapeutic strategies is still lacking. This study aimed to develop a novel gene signature related to anti-androgen resistance and evaluate its prognostic and therapeutic implications. METHODS: Anti-androgen resistance-related differentially expressed&#xa0;genes (ARRDEGs) were identified through transcriptomic analysis of enzalutamide- and dual enzalutamide abiraterone-resistant PCa cell lines from the GEO database. Functional enrichment analysis was performed to determine the biological roles of these genes. A prognostic gene signature was developed using univariate Cox regression, LASSO, and multivariate Cox regression models. The model was validated in independent PCa cohorts from The Cancer Genome Atlas (TCGA). Additionally, we assessed the correlation between the signature, immune infiltration, immune checkpoint expression, and drug sensitivity. The efficacy of PLK1 inhibition combined with enzalutamide was further explored using in vitro and in vivo experiments. RESULTS: We identified 304 ARRDEGs, from which three key genes (LMNB1, SSPO, and PLK1) were selected to construct a prognostic signature. This gene signature effectively stratified PCa patients into high- and low-risk groups, with the high-risk group exhibiting shorter recurrence-free survival and distinct immune characteristics. High-risk patients demonstrated elevated immune checkpoint expression (B7H3, CTLA-4, B7-1, and TIGIT), increased M2 macrophage infiltration, and enhanced sensitivity to chemotherapy and targeted therapy. Mechanistically, PLK1 inhibition potentiated the antitumor effect of enzalutamide by downregulating SLC7A11 and inducing ferroptosis, providing a potential therapeutic strategy to overcome anti-androgen resistance. CONCLUSION: We established a novel ARRDEGs-based prognostic signature that predicts PCa progression and response to chemotherapy&#xa0;and targeted therapy. The integration of this signature with immune profiling and drug sensitivity analysis provides a valuable tool for precision oncology in PCa. Our findings highlight the potential of PLK1 inhibition as a therapeutic strategy to enhance enzalutamide efficacy and overcome resistance.

Humans

Liver transcriptome analysis revealed multiple immune processes and lipid metabolism pathways involved in the defense response of the turbot (Scophthalmus maximus) against Aeromonas salmonicida.

Aeromonas salmonicida is a significant pathogen causing notable economic losses in Scophthalmus maximus aquaculture. This study utilized Illumina sequencing technology to examine the transcriptional response characteristics of S. maximus liver at 24&#xa0;h following A. salmonicida infection. A total of 2363 differentially expressed genes (DEGs) were identified when compared to the negative control group. The immunity-related Toll-like receptor signaling pathway, NOD-like receptor signaling pathway, as well as metabolism-related PPAR signaling pathway and insulin signaling pathway, were notably enriched. Significant differences exist in the expression of key genes within the PPAR pathway, particularly cd36, acsl4a, ppar&#x3b1;a, and plin2, all of which mediate the interaction between lipid metabolism and the immune response. These results offer valuable insights into the immunometabolic regulatory mechanism of S. maximus response to A. salmonicida infection.

Animals

Histology-Based Virtual RNA Inference Identifies Pathways Associated With Metastasis Risk in Colorectal Cancer.

Colorectal cancer (CRC) remains a major health concern, with >150,000 new diagnoses and >50,000 deaths annually in the United States, underscoring an urgent need for improved screening, prognostication, disease management, and therapeutic approaches. The tumor microenvironment (TME)-comprising cancerous and immune cells interacting within the tumor's spatial architecture-plays a critical role in disease progression and treatment outcomes, reinforcing its importance as a prognostic marker for metastasis and recurrence risk. However, traditional methods for TME characterization, such as bulk transcriptomics and multiplex protein assays, lack sufficient spatial resolution. Although spatial transcriptomics (ST) allows for the high-resolution mapping of whole transcriptomes at near-cellular resolution, current ST technologies (eg, Visium and Xenium) are limited by high costs, low throughput, and issues with reproducibility, preventing their widespread application in large-scale molecular epidemiology studies. In this study, we refined and implemented virtual RNA inference (VRI) to derive ST-level molecular information directly from hematoxylin and eosin (H&E)-stained tissue images. Our VRI models were trained on the largest matched CRC ST data set to date, comprising 45 patients and >300,000 Visium spots from primary tumors. Using state-of-the-art deep learning models (UNI, ResNet-50, Vision Transformer, and Vision Mamba), we achieved a median Spearman's correlation coefficient of 0.546 between predicted and measured spot-level expression. As validation, VRI-derived gene signatures linked to specific tissue regions (tumor, interface, submucosa, stroma, serosa, muscularis, and inflammation) showed strong concordance with signatures generated via direct ST, and VRI performed accurately in estimating cell-type proportions spatially from H&E slides. In an expanded CRC cohort controlling for tumor invasiveness and clinical factors, we further identified VRI-derived gene signatures significantly associated with key prognostic outcomes, including metastasis status. Although certain tumor-related pathways are not fully captured by histology alone, our findings highlight the ability of VRI to infer a wide range of "histology-associated" biological pathways at near-cellular resolution without requiring ST profiling. Future efforts will extend this framework to expand TME phenotyping from standard H&E tissue images, with the potential to accelerate translational CRC research at scale.

Humans