Search PubMedSearch

SEARCH · Search PubMed

Results for “transcriptomic analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Comprehensive Transcriptome Annotation of Thousands of HIV-1 Genomes.

Alternative splicing in HIV-1 has been a central focus of decades of research, uncovering key mechanisms of viral gene regulation, immune evasion, and therapeutic response - yet, no reference resource has existed to support transcriptome-wide analysis, limiting adoption of modern computational methods. We present HIV Atlas (https://ccb.jhu.edu/HIV_Atlas), the first reference-quality annotation of HIV-1 and SIV transcriptional diversity. We manually curated transcriptomes for HIV-1HXB2 and SIVmac239 and developed Vira, an automated annotation-transfer method specifically designed to address unique challenges of viral genome biology, to generate high-quality annotations for 2,077 complete HIV-1 genomes. Using the resources presented in our work, we evaluated conservation of splice sites, revealing near-perfect preservation of major donors and acceptors. Furthermore, using several public datasets, we demonstrate how HIV Atlas enhances methodology, improves the quality and novelty of results, and opens novel avenues for research, supporting more accurate and comprehensive analyses of bulk, single-cell, and spatial RNA-seq in HIV-1 studies.

Journal Article

Genetic association between epilepsy and gliomas: Insights from Mendelian randomization and single-cell transcriptomic analyses.

BACKGROUND: Seizures are prevalent in glioma patients, especially in those with low-grade gliomas. The interaction between gliomas and epilepsy involves complex biological mechanisms that are not fully understood. METHODS: We collected Genome-Wide Association Study data for epilepsy and gliomas, performed differential expression analysis, and conducted Gene Ontology (GO) enrichment analysis on the identified genes. Single-cell RNA sequencing data (scRNA-seq) from GSE221534 dataset in Gene Expression Omnibus (GEO) were used to analyze cell-cell interactions within glioma samples from patients with and without epilepsy. RESULTS: Mendelian Randomization (MR) analysis revealed significant associations between genetic variants related to epilepsy and glioma risk, suggesting a potential causal relationship, especially in astrocytomas. Differential expression analysis identified epilepsy-related genes that were significantly upregulated in astrocytoma tissues compared to normal brain tissues. GO enrichment analysis indicated that these genes are involved in critical biological processes such as neurogenesis and cellular signaling. The scRNA-seq analysis showed, compared to non-epileptic samples, glioma stem cells, microglia, and NK cells are increased in the core regions of astrocytomas in epileptic patients. Additionally, intercellular communication between tumor cells and other non-tumor cells is markedly enhanced in astrocytoma samples from epileptic patients. CONCLUSION: This study provides evidence of a genetic association between epilepsy and gliomas and elucidates the biological mechanisms through which epilepsy may influence glioma progression.

Humans

Unsupervised multiscale clustering of single-cell transcriptomes to identify hierarchical structures of cell subtypes.

BACKGROUND: Cell clustering is an essential step in uncovering cellular architectures in single-cell RNA sequencing (scRNA-seq) data. However, the existing cell clustering approaches are not well designed to dissect complex structures of cellular landscapes at a finer resolution. RESULTS: Here, we develop a multiscale clustering (MSC) approach to construct a sparse cell-cell correlation network for unsupervised identification of de novo cell types and subtypes across multiple resolutions. Based upon simulated silver- and gold-standard data as well as real scRNA-seq data in diseases, MSC demonstrates significantly improved performance compared to established benchmark methods and reveals a biologically meaningful cell hierarchy to facilitate the discovery of novel disease-associated cell subtypes and mechanisms. CONCLUSIONS: We present MSC as a new single-cell multiscale clustering framework as a powerful tool for advancing discoveries in disease-associated cell populations using single-cell sequencing data.

Single-Cell Analysis

RNF43 Mutations Are Associated With the Classical Molecular Subtype, Vigorous Antitumor Immune Responses, and Prolonged Survival in Pancreatic Adenocarcinoma.

RNF43 mutations were correlated with microsatellite status in colorectal cancer and with fewer and later recurrences in pancreatic ductal adenocarcinoma (PDAC). Here, we undertake a detailed assessment of RNF43 mutations in PDAC. A total of 313 PDACs (308 microsatellite stable [MSS] and 5 microsatellite-instable [MSI] cases) underwent next-generation sequencing (Oncomine Tumor Mutation Load assay; Thermo Fisher). Spatial analyses (NanoString) classified PDACs according to their transcriptomic and proteomic immune signaling. Fluorescent imaging was used to define spatial compartments (tumor: pancytokeratin+/CD45- and leukocytes: pancytokeratin-/CD45+). Each of 20 PDACs with RNF43 mutations (RNF43mut) and without RNF43 mutations (RNF43wt) underwent multiplex immunofluorescence analysis to determine immune status. A total of 153 PDACs (22 RNF43mut and 131 RNF43wt cases) underwent bulk RNA sequencing to assign into molecular subtypes. Overall, 24 RNF43 mutations were identified (22 MSS PDACs and 2 MSI PDACs). The incidence of RNF43 mutations in MSS PDACs (7.1%) was consistent with The Cancer Genome Atlas (6.7%). However, RNF43 mutations were more frequent among MSI PDACs (40%). Additionally, RNF43mut had differential frequencies of other mutations (including Wnt pathway genes), higher tumor mutational burden values (5.5 mut/mb vs 1.67 mut/mb; P < .01), and significantly longer overall survival (47 vs 18 months; P < .0001) than RNF43wt. Moreover, RNF43mut exhibited significantly higher densities of CD8+ T lymphocytes, dendritic cells, and B lymphocytes (P < .001) and an upregulation of ITGAX, CD11c, CD8, and HLA-DR compared with RNF43wt. Patients with RNF43mut PDACs were more often of the classical molecular subtype (20/22, 90.9%). RNF43mut PDACs showed high tumor mutational burden values, suggesting increased neoantigen load coupled with an abundance of antigen-presenting immune cells and an upregulation of immune determinants promoting antigen presentation. All this contributes to stronger antitumor immune responses and improved clinical outcomes.

Humans

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning

Genome-wide identification and cold-stress-responsive expression analysis of the NOX gene family in Cucumis melo.

NADPH oxidases (NOXs) are crucial enzymes for reactive oxygen species (ROS) generation in plants and play vital roles in growth, development, and stress responses. To elucidate the sequence characteristics of the NOX gene family and its low-temperature response patterns in melon (Cucumis melo L.), this study conducted genome-wide identification and expression profiling of NOX family members using bioinformatics analysis, RNA-seq transcriptome sequencing, and real-time quantitative PCR (RT-qPCR). The results revealed that eight NOX members were identified in the melon genome, distributed across six chromosomes. All members harbored conserved domains including Ferric_reductase, FAD_binding_8, NAD_binding_6, and NADPH_Ox, and the encoded proteins were generally basic and hydrophilic. Phylogenetic analysis classified the NOX proteins into five subgroups. Synteny analysis indicated the presence of only one pair of intraspecific duplicated genes in melon, which was under purifying selection. The promoter regions contained multiple hormone- and stress-responsive cis-acting elements, with CmNOX2 and CmNOX4 harboring low-temperature responsive elements. Following treatment at 4&#x2103; for 24 h and 48 h, leaf relative electrolyte leakage (REL) increased from 28.33% to 42.67% and 52.67%, respectively; transcriptome analysis identified 5,633 and 6,882 differentially expressed genes (DEGs), respectively. Cold-responsive genes exhibited significant differential expression, with SLAC1 and CPK19 showing sustained upregulation. RT-qPCR results demonstrated that the expression of CmNOX2, CmNOX5, CmNOX6, and CmNOX7 was significantly downregulated after low-temperature treatment, whereas CmNOX4 expression was significantly upregulated at 48 h. Integrating promoter elements and expression characteristics, CmNOX4 may represent an important candidate gene involved in melon low-temperature response. This study systematically characterized the structure, evolution, and expression patterns of the melon NOX gene family, identified candidate genes responsive to low temperature, and provides a reference for further investigation into the mechanisms underlying melon cold adaptation.

Cucumis melo

Comparative analysis of DREB gene family in buckwheat: the role of FtDREB02 in the delphinidin biosynthesis and drought stress response.

Dehydration response element binding (DREB) transcription factors play a pivotal role in plant abiotic stress responses, but its evolutionary and functional characterization in buckwheat remains unexplored. Here, we conducted a comprehensive analysis of the DREB gene family across three buckwheat species, revealing segmental duplication as the primary driver of family expansion and potential purifying selection during evolution. A FtDREB02 gene, classified as group A2, was identified through genome-wide association analysis (GWAS) on drought tolerance and delphinidin content. Functional validation in Arabidopsis thaliana and the hairy root of Tartary buckwheat (Fagopyrum tataricum) demonstrated that overexpression of this gene promotes delphinidin biosynthesis and enhances plant resistance to water scarcity. Through the integration of DAP-seq and PEG transcriptome cluster analysis, a FtANS candidate was screened. Functional studies showed that FtDREB02 regulates delphinidin content by binding directly to DRE elements of the FtANS promoter. This research identifies and comprehensively analyzes the DREB family within buckwheat species, elucidating the regulatory mechanisms of FtDREB02 in controlling flavonoid biosynthesis and drought resistance, providing potential genetic resources for breeding buckwheat varieties with excellent agronomic traits.

Anthocyanins

Single-cell transcriptomics reveals heterogeneous stress responses and Mg2+-mediated survival mechanisms in Lactobacillus delbrueckii subsp. bulgaricus during freeze-drying and storage.

Maintaining the viability of lactic acid bacteria during dehydration and subsequent storage remains a significant challenge. Here, we employed single-cell RNA sequencing to reveal the heterogeneous stress responses of Lactobacillus delbrueckii subsp. bulgaricus, identifying seven distinct transcriptional clusters across the liquid culture, freeze-drying, and storage phases. The dominant clusters in the freeze-drying and storage were not completely consistent, showing significant functional differentiation. Genomic stability may be important for survival during freeze-drying and storage, while intracellular energy homeostasis appears important for viability during storage. The magnesium transporter mgtB was highly expressed in clusters tolerant to freeze-drying and storage, suggesting a critical role for Mg2+ homeostasis. Further experimental validation confirmed that Mg2+ treatment significantly bolstered stress resistance, increasing immediate post-freeze-drying survival by over 2-fold (up to 92.90%) and post-storage survival by over 5-fold (up to 5.98%). Proteomic data indicated that Mg2+ supplementation correlated with the maintenance of several biological functions potentially relevant to bacterial survival during freeze-drying and storage, including DNA repair, translation, and central carbon metabolism. These findings provide a map of microbial stress resistance through population heterogeneity and offer a potential strategy that may be adapted for enhancing the stability of other industrial lactic acid bacteria products.

Freeze Drying

Multilevel genomic, transcriptomic, and epidemiologic evidence linking diabetic retinopathy to Alzheimer disease.

BACKGROUND: Diabetic retinopathy (DR) and Alzheimer disease (AD) share metabolic and vascular dysfunctions, but the extent to which they reflect overlapping genetic susceptibility and neurovascular-metabolic regulatory pathways remains unclear. We combined multi-omics analyses with population-based data to examine the genetic convergence, cellular pathways, and longitudinal association between DR and AD. METHODS: We performed a two-sample Mendelian randomisation (MR) to estimate the association between genetically predicted DR liability and AD risk. We used Bayesian colocalisation analysis to identify shared genomic loci, and summary-data-based MR (SMR) to detect expression-mediated genes jointly associated with DR and AD. We analysed single-cell RNA sequencing data to characterise shared cellular features and related biological pathways. We also conducted an MR-based mediation analysis to explore whether lipid-related, metabolic, or inflammatory traits mediated the observed DR-AD association, and a longitudinal analysis of the UK Biobank cohort to assess the association between DR and incident AD. RESULTS: With the MR analysis, we found that genetically predicted liability to DR was associated with a modest increase in AD risk. Colocalisation analysis supported a shared genetic signal. We identified three genes with shared expression-mediated associations across DR and AD through SMR. Functional enrichment analyses revealed partially overlapping neurovascular and metabolic pathways. Using MR-based mediation analysis, we found no significant intermediary traits linking DR and AD. Findings from the UK Biobank cohort were directionally consistent with the genetic analyses. CONCLUSIONS: Genetic liability to DR is associated with an increased risk of AD and is accompanied by shared expression-mediated effects and convergent neurovascular-metabolic pathways. These findings support the possibility that DR may serve as a clinically accessible indicator of increased neurodegenerative vulnerability.

Humans

Repeated emergence and fitness heterogeneity of KPC-33 in ST11 Klebsiella pneumoniae under ceftazidime-avibactam pressure.

Ceftazidime-avibactam (CZA) is an important therapeutic option for infections caused by Klebsiella pneumoniae carbapenemase (KPC)-producing Klebsiella pneumoniae. However, CZA exposure also selects for emergent KPC variants. Their in vivo evolutionary patterns, fitness consequences, and underlying molecular mechanisms remain unclear. We performed a longitudinal multiomics analysis of 35 clonally related ST11 KPC-producing K. pneumoniae isolates collected from eight hospitalized patients during clinical follow-up, most of whom had received CZA therapy. Whole-genome sequencing, antimicrobial susceptibility testing, in vitro competition assays, enzyme kinetic analysis, and transcriptomic sequencing were used to systematically characterize the within-host evolutionary dynamics of KPC variants and the fitness heterogeneity of KPC-33. Multiple KPC variants were identified during longitudinal follow-up, among which KPC-33 was the most frequently detected. Among the seven patients who received CZA treatment, KPC-33 was detected in longitudinal isolates from four patients. It was also identified in patient P3, who had not received CZA, whereas other variants were only sporadically identified. Biochemical analysis showed that KPC-33 exhibited an altered kinetic profile relative to KPC-2, characterized by reduced catalytic turnover and altered substrate affinity. KPC-33 did not exhibit a uniform and pronounced fitness defect but instead showed marked strain-dependent heterogeneity. Strains with higher competitive fitness generally showed only limited transcriptional changes, whereas those with lower fitness were accompanied by broader transcriptional remodeling. In this longitudinal cohort, KPC-33 was repeatedly detected, predominantly under CZA-associated selective conditions. Its fitness consequences were clearly strain background dependent and may be associated with the extent of transcriptional remodeling. These findings provide new evidence for understanding the in vivo evolution of CZA resistance.

KPC-33

PLAID: ultrafast single-sample gene set enrichment scoring.

SUMMARY: In recent years, computational methods have emerged that calculate enrichment of gene signatures within individual samples. These signatures offer critical insights into the coordinated activity of functionally related genes, proteins or metabolites, enabling the identification of unique molecular profiles in individual cells and patients. This strategy is pivotal for patient stratification and advancement of personalized medicine. However, the rise of large-scale datasets, including single-cell profiles and population biobanks, has exposed significant computational inefficiencies in existing methods. Current methods often demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming these limitations is a focus of current efforts by bioinformatics teams in academia and the pharmaceutical industry, as essential to support basic and clinical biomedical research. To address this critical need, we developed PLAID (Pathway Level Average Intensity Detection), an ultrafast and memory optimized single sample gene set enrichment algorithm that utilizes sparse matrix computation. PLAID delivers highly accurate gene set scoring and surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data. PLAID uniquely integrates the most widely used gene set scoring algorithms, enabling researchers to apply multiple methods for cross-validation with outstanding runtime efficiency and minimal memory requirement. AVAILABILITY AND IMPLEMENTATION: PLAID is implemented in the R language for statistical computing. PLAID source code and installation instructions are available with no restrictions at https://github.com/bigomics/plaid.

Algorithms

Blood phenylalanine lowering partially reverses white matter changes in a mouse model of phenylketonuria.

Phenylketonuria (PKU) is a genetic defect caused by lack of the liver enzyme phenylalanine hydroxylase (PAH). This deficiency results in elevated blood phenylalanine (Phe) levels and neurotoxicity, which is manifested by reduced brain size, lower neurotransmitter levels, and reduced myelination. The goal of this study was to investigate brain myelination defects and their reversibility upon blood Phe lowering by analyzing the corpus callosum (CC) of adult Pahenu2 (PAH-deficient) mice. MRI and immunostaining demonstrated a significant reduction in CC volume in Pahenu2 mice. Treatment with an adeno-associated vector (AAV) encoding mouse PAH for 3.5 months improved but did not completely normalize CC volume. Total cholesterol, a major component of myelin, was unchanged in the CC of Pahenu2 mouse, while some sterol intermediates were significantly reduced by treatment. Single-nuclei transcriptomics showed an upregulation of oxidative stress-related pathways and increased expression of transthyretin, ApoE, Cst3, and Cd81 in CC in Pahenu2 mice. Normalization of blood Phe restored gene expression to levels comparable to those of heterozygous mice and was associated with the generation of differentiated myelin-producing oligodendrocyte subtypes and neuroprotective astrocytes. In summary, Pahenu2 mice showed white matter abnormalities and changes in transcriptome and sterol profiles, which were partially corrected by the normalization of blood Phe.

Animals

Hepatic metabolic adaptation to endurance exercise: temporal and sex differences by multiomics integration and validation.

BACKGROUND: Although endurance exercise benefits liver health, sex-specific adaptive trajectories remain unclear. This study mapped dynamic liver adaptation in males and females during prolonged training and identified underlying molecular programs. METHODS: Using publicly available time-resolved liver multi-omics data generated by the Molecular Transducers of Physical Activity Consortium (MoTrPAC), we established a computational pipeline for differential analysis of transcriptomic, proteomic, phosphoproteomic, and metabolomic data with FDR correction, followed by FGSEA pathway enrichment. Kinase activities were inferred through ortholog mapping and PhosphoSitePlus. Cross-omics co-expression networks were constructed using WGCNA and topological overlap to link omics features with physiological phenotypes. For experimental validation, liver tissues were collected from endurance-trained Sprague-Dawley rats, and key nodes were confirmed by Western blotting, qRT-PCR, and immunofluorescence/immunohistochemical staining. Public scRNA-seq data were further integrated to map multi-omics signals to single-cell resolution and assess functional changes in specific cell types. RESULTS: The hepatic response to exercise stress was stage-specific, shifting from early transcriptional activation to later proteomic and metabolic remodeling. Multi-omics integration revealed distinct sex-associated adaptive trajectories: males were more strongly associated with energy metabolism, redox-related programs, and amino acid/organic acid catabolism, whereas females showed prominent membrane lipid remodeling, proteostasis -related programs, and mitochondrial/ribosomal translational features. Single-cell analysis showed that tissue remodeling occurred without major lineage turnover, instead involving altered communication among pre-existing cell communities. Validation of PPP1R3G identified a protein-dominant exercise-responsive marker, supporting the contribution of post-transcriptional or protein-level regulation. CONCLUSIONS: Hepatic adaptation to endurance stress follows a cross-omics evolutionary pattern with sex-specific reprogramming of energy supply and homeostatic maintenance. This time-resolved framework clarifies how exercise improves liver function and supports sex-oriented metabolic interventions and therapeutic target discovery.

Animals

Multi-omics technologies: Novel tools and methods for assessing nerve injury and regeneration.

Recently, with the rapid advancement of multi-omics technologies, including genomics, transcriptomics, proteomics, and metabolomics, new tools and approaches have been introduced for studying nerve injury and regeneration. This review highlights the application and progress of multi-omics in uncovering the mechanisms of nerve injury, guiding the development of regenerative strategies, and promoting clinical translation. By integrating multi-omics datasets, researchers can comprehensively track dynamic molecular changes following nerve injury, including abnormal gene expression, disrupted protein signaling, altered metabolic programs, and shifts in the immune microenvironment. Single-cell multi-omics technologies resolve cellular heterogeneity, revealing the distinct functions of neurons, glial cells, and immune cell subpopulations during the injury response. Spatially resolved transcriptomics maintain the spatial context of lesion and regeneration sites, enabling precise localization for targeted interventions. Multi-omics technologies not only identify key molecular players involved in nerve regeneration but also create opportunities for personalized medicine. Nonetheless, integrating multi-omics data poses technical challenges, including high dimensionality, batch effects, and algorithmic constraints, while ethical concerns related to stem cell therapy and gene editing require stringent oversight. To transition from structural reconstruction to functional remodeling, future research should emphasize artificial intelligence-driven data integration, organ-on-a-chip modeling, and cross-disciplinary collaboration to overcome existing technical barriers and accelerate the clinical application of neuroregenerative therapies.

artificial intelligence

Coalescing single-cell genomes and transcriptomes to decode breast cancer progression.

Understanding epithelial lineages of breast cancer and genotype-phenotype relationships requires direct measurements of the genome and transcriptome of the same single cells at scale. To achieve this, we developed wellDR-seq, a high-genomic-resolution, high-throughput method to simultaneously profile the genome and transcriptome of thousands of single cells. We profiled 33,646 single cells from 12 estrogen-receptor-positive breast cancers and identified ancestral subclones in multiple patients that showed a luminal hormone-responsive lineage, indicating a potential cell of origin. In contrast to bulk studies, wellDR-seq enabled the study of subclone-level gene-dosage relationships, which showed near-linear correlations in large chromosomal segments and extensive variation at the single-gene level. We identified dosage-sensitive and dosage-insensitive genes, including many breast cancer genes as well as sporadic copy-number aberrations in non-cancer cells. Overall, these data reveal complex relationships between copy number and gene expression in single cells, improving our understanding of breast cancer progression.

Breast Neoplasms

Deconstructing the Alternative Lengthening of Telomeres: Integromics Prioritizes Five Master Hubs Dictating Clinical Survival and Therapeutic Vulnerabilities.

The Alternative Lengthening of Telomeres (ALT) pathway drives replicative immortality in aggressive malignancies, particularly sarcomas and gliomas. Clinical ALT stratification has relied on screening for structural ATRX and DAXX mutations. However, this genotypic approach fails to capture the dynamic macro-reprogramming required to sustain ALT. Here, we established and validated a 28-gene transcriptomic signature that captures the ALT-associated transcriptomic phenotype of the ALT phenotype. Using multivariate Cox proportional hazards models and time-dependent ROC analyses, we demonstrate that this signature is a robust, independent predictor of poor overall survival in Sarcoma (SARC) and Lower Grade Glioma (LGG) cohorts, outperforming the prognostic value of traditional ATRX/DAXX mutational status. Genomic mapping revealed this transcriptional synchrony is structurally facilitated by non-random focal clustering on Chromosome 8. To deconstruct the machinery driving this lethal phenotype, we employed an integromic approach, synthesizing protein-protein and metabolic flux networks. Topological algorithms prioritized five indispensable hubs: TP53, ATM, ATR, PCNA, and UBE2I. Gene-metabolite profiling identified PCNA as a bottleneck funneling extreme deoxyribonucleotide (dNTP) demand to sustain break-induced telomeric recombination. To translate these vulnerabilities into actionable treatments, we mapped these hubs to a precision pharmacological network. We propose a multi-targeted strategy combining FDA-approved PARP inhibitors to exploit ATR-mediated synthetic lethality, alongside antimetabolites to induce nucleotide starvation. This study redefines ALT risk stratification and provides a data-driven framework to target and treat resistant ALT-positive tumors.

Alternative Lengthening of Telomeres

An integrated single-cell and spatial transcriptomic atlas of thyroid cancer progression identifies prognostic fibroblast subpopulations.

Although well-differentiated thyroid carcinoma (WDTC) is characterized by a robust treatment response, aggressive subtypes, such as anaplastic thyroid carcinoma (ATC), remain highly lethal. To understand thyroid cancer evolution in both children and adults, we analyzed single-cell transcriptomes of 423,733 cells from 81 samples and spatially resolved key tumor and microenvironment populations across 28 tumors with spatial transcriptomics, including rare and unique composite WDTC/ATC tumors and pediatric diffuse sclerosing thyroid carcinomas. Additionally, we identified gene signatures of stromal cell populations in 5 large thyroid cancer bulk RNA-sequencing cohorts. Through this multi-institutional effort, we defined a population of POSTN+ myofibroblast cancer-associated fibroblasts (myCAFs) that are intimately associated with invasive tumor cells and correlate with poor prognosis, lymph node metastasis, and disease progression in thyroid carcinoma. We also revealed a population of inflammatory CAFs that are distant to tumor cells and are found in the inflammatory stromal microenvironment of autoimmune thyroiditis. Together, our study provides spatial profiling of thyroid cancer evolution in samples with mixed WDTC/ATC histopathology and identifies a prognostic myCAF subtype with potential clinical utility in predicting aggressive disease in both children and adults.

Humans

A multi-modal survival prediction framework with group-based batch training and structural consistency alignment.

OBJECTIVE: Integrating whole-slide images (WSIs) with transcriptomic profiles is pivotal for enhancing cancer survival prediction. However, the intrinsic gigapixel resolution and variable sequence lengths of WSIs create a fundamental trade-off between training efficiency and the preservation of data heterogeneity in existing frameworks. Furthermore, substantial statistical and structural discrepancies between histological and genomic modalities often impede effective cross-modal alignment and fusion, thereby limiting prognostic accuracy. METHODS: We propose PRISM, an efficient multi-modal learning framework for integrating WSIs with transcriptomic profiles. To reconcile training efficiency with full data heterogeneity, PRISM first stochastically partitions variable-length WSI sequences into a main subset and a complementary residual subset, both of which are packed into fixed-length groups for batch training. The main subset is processed in the main branch, utilizing isolation masking to maintain intra-group sequence independence. Simultaneously, the residual subset is consolidated into "hyperslides" within a residual branch that leverages tailored supervision, effectively capturing inter-slide correlations. Furthermore, PRISM integrates an Informative Token Aggregation (ITA) module to reduce redundancy in WSIs and employs Cross-batch Structural Consistency Alignment (CBSCA) mechanism to enhance inter-modal structural connectivity. Finally, efficient cross-modal feature interaction is achieved through a Low-rank Bilinear Gated Fusion (LBGF) module. Code is available at https://github.com/Alisa2080/PRISM. RESULTS: Compared with existing methods, PRISM achieves the best overall C-index across five TCGA cohorts. On the larger TCGA-BRCA dataset, PRISM requires only 6&#xa0;hours of training time, substantially reducing computational cost relative to strong multimodal baselines. Furthermore, comprehensive evaluations demonstrate that PRISM achieves the best overall IBS ranking and favorable time-dependent AUC performance at 1, 3, and 5&#xa0;years, thereby delivering a more favorable trade-off between prognostic performance and computational efficiency. CONCLUSION: PRISM provides a favorable balance between predictive performance, calibration quality, and computational efficiency, highlighting its potential for practical deployment in multimodal survival modeling for computational pathology.

Humans