Search PubMedSearch

SEARCH · Search PubMed

Results for “transcriptome annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The chromosome-level genome of Stylosanthes guianensis provides insights into genome evolution and environmental adaptation.

Stylosanthes guianensis is a leguminous forage crop of significant economic importance, primarily distributed in tropical and subtropical regions. It exhibits strong adaptability to various stresses, yet the genetic basis underlying this trait remains unclear. In this study, we constructed the first chromosome-scale reference genome of S. guianensis using a combination of Nanopore and Hi-C sequencing technologies. The assembled genome size is 1254 Mb, with 10 pseudochromosomes. Using Nanopore full-length transcriptome data, we generated high-quality transcript-level gene annotations, identifying 36 585 gene models and 110 601 transcripts. The repetitive sequences in S. guianensis account for 79.16% of the genome, with the extensive expansion of Gypsy elements in long terminal repeats contributing to its genome size enlargement. Comparative genomic and transcriptomic analyses revealed that flavonoid metabolism plays a pivotal role in stress adaptation, providing new insights into the genetic basis of stress tolerance. Additionally, we generated whole-genome methylation profiles under cold treatment and control conditions, offering valuable data for future epigenomic research. These findings provide essential molecular resources for understanding stress resilience in S. guianensis and advancing its molecular breeding.

Genome, Plant

Whole blood transcriptome profile identifies motor neurone disease RNA biomarker signatures.

Blood-based biomarkers for motor neuron disease are needed for better diagnosis, progression prediction, and clinical trial monitoring. We used whole blood-derived total RNA and performed whole transcriptome analysis to compare the gene expression profiles in (motor neurone disease) MND patients to the control subjects. We compared 42 MND patients to 42 aged and sex-matched healthy controls and described the whole transcriptome profile characteristic for MND. In addition to the formal differential analysis, we performed functional annotation of the genomics data and identified the molecular pathways that are differentially regulated in MND patients. We identified 12,972 genes differentially expressed in the blood of MND patients compared to age and sex-matched controls. Functional genomic annotation identified activation of the pathways related to neurodegeneration, RNA transcription, RNA splicing and extracellular matrix reorganisation. Blood-based whole transcriptomic analysis can reliably differentiate MND patients from controls and can provide useful information for the clinical management of the disease and clinical trials.

Humans

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics

De Novo Genome Sequence Assembly of the Algal Endosymbiont Micractinium conductrix Derived From Its Host Paramecium bursaria 186b.

Endosymbiosis is a major driver of evolutionary innovation and underpins the function of diverse ecosystems. The origins and evolution of endosymbiosis are challenging to study experimentally due to the short-lived culturability of many microbial strains derived from endosymbiotic interactions. The facultative endosymbiosis between the ciliate, Paramecium bursaria, and the green alga, Micractinium conductrix (Chlorellaceae, Trebouxiophyceae), is ecologically widespread and has emerged as a powerful lab-tractable model system. This endosymbiosis is founded upon a reciprocal nutrient exchange, but each of the species can be cultured independently enabling quantification of symbiotic fitness effects, new partnerships to be generated in the lab, and co-associations to be subject to experimental evolution. To date, evolve-and-resequence approaches have been limited due to a lack of high-quality genome assemblies enabling gene variants to be identified. Here, we report a near telomere-to-telomere genome assembly for M. conductrix 186b, using a range of sequencing technologies. Comparative analysis shows that this is one of the most complete Chlorellaceae algal genome assemblies available to date. To aid accurate gene calling and annotation, we conducted both RNAseq and Iso-Seq transcriptome sequencing experiments. Collectively, these 'omics datasets will facilitate: (i) comparative genomics studies of endosymbiont evolution, (ii) evolve-and-resequence experiments, (iii) genome-scale metabolic modeling studies, and (iv) identification of targets for genetic modification experiments and biotechnological applications.

Symbiosis

Spatial genomics: Mapping the landscape of fibrosis.

Organ fibrosis causes major morbidity and mortality worldwide. Treatments for fibrosis are limited, with organ transplantation being the only cure. Here, we review how various state-of-the-art spatial genomics approaches are being deployed to interrogate fibrosis across multiple organs, providing exciting insights into fibrotic disease pathogenesis. These include the detailed topographical annotation of pathogenic cell populations and states, detection of transcriptomic perturbations in morphologically normal tissue, characterization of fibrotic and homeostatic niches and their cellular constituents, and in situ interrogation of ligand-receptor interactions within these microenvironments. Together, these powerful readouts enable detailed analysis of fibrosis evolution across time and space.

Humans

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59 Mb in 31 scaffolds with an N50 length of 33.98 Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals

Transcriptomic responses of Porphyrophora sophorae larvae during licorice root colonization reveal coordinated remodeling of translation, mitochondrial energy metabolism and defense-related genes.

BACKGROUND: Porphyrophora sophorae is a subterranean piercing-sucking scale insect that damages licorice (Glycyrrhiza uralensis) roots, but the molecular responses associated with larval root colonization remain insufficiently defined. METHODS: We compared non-parasitic larvae (NP) and root-colonizing larvae (RC) using six RNA-seq libraries, de novo transcriptome assembly, DESeq2-based differential expression analysis, GO/KEGG enrichment, annotation-based candidate gene screening, and RT-qPCR validation of selected genes. RESULTS: Sequencing yielded 260.91 million clean reads, and de novo assembly produced 60,794 non-redundant transcripts. DESeq2 identified 703 FDR-significant DEGs, including 49 upregulated and 654 downregulated genes in RC larvae. Upregulated genes were mainly associated with translation- and ribosome-related processes, whereas downregulated genes were enriched in mitochondrial, oxidation-reduction, energy metabolism, and oxidative phosphorylation-related functions. Annotation-based screening identified 75 FDR-significant candidate genes associated with chemosensation, defense-related responses, and energy metabolism, with mitochondrial energy metabolism-related genes forming the largest module. RT-qPCR validation based on the raw Ct data showed concordant expression directions for ten selected transcript targets. CONCLUSIONS: Root colonization in P. sophorae larvae was associated with coordinated transcriptional remodeling involving selective activation of translation-related processes, adjustment of mitochondrial energy metabolism, and changes in defense-related gene expression. These results provide candidate molecular targets for future functional studies of host contact, feeding establishment, and physiological adjustment in this subterranean scale insect.

Animals

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 × 50 DNA nanoballs and an approximate nominal footprint of 25 × 25 µm, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal‑Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

Systematic discovery of retina-enriched Rik genes identifies 1190005I06Rik as a novel modulator of visual signalling.

BACKGROUND: High‑throughput transcriptome projects have revealed thousands of mammalian genes with little or no functional annotation. Among these are hundreds of loci assigned provisional “Rik” identifiers following discovery in the RIKEN cDNA annotation effort. Although often dismissed as genomic dark matter, such genes may encode tissue‑restricted proteins that modulate physiologic functions and influence disease. The retina is a highly specialised neural tissue and a common site of inherited disorders; understanding its molecular repertoire could illuminate novel therapeutic avenues. METHODS: We integrated bulk RNA‑seq from ten adult mouse tissues, evolutionary and domain analysis, single‑cell RNA‑seq, and CRISPR/Cas9 gene disruption to systematically catalogue protein‑coding Rik genes enriched in the retina and test the function of a representative gene. RESULTS: A rigorous differential expression analysis identified 44 Rik genes with robust retina‑specific expression compared with nine non‑retinal tissues. Many of these genes lack orthologues beyond rodents, while others show broad conservation, illustrating a continuum from lineage‑restricted to conserved retinopathy candidates. Single‑cell transcriptomics revealed that these genes are expressed across retinal cell types, with the highest aggregate expression in cone photoreceptors and inner interneurons. To evaluate physiological significance, we generated a 1190005I06Rik knockout mouse. Although retinal architecture appeared normal, loss of 1190005I06Rik enhanced electroretinogram b‑wave amplitudes and altered light‑avoidance behaviour, indicating that this previously uncharacterised gene acts as a negative modulator of visual signalling. CONCLUSIONS: We present a curated atlas of retina‑enriched Rik genes and demonstrate that 1190005I06RIK modulates retinal circuit function. This resource expands the molecular landscape of the retina and provides new candidates for the genetic basis of inherited retinal disease. Our findings underscore that unannotated genes may exert measurable effects on sensory processing and warrant systematic exploration in the context of human ocular disorders.

Animals

Post-genome-wide association study dissects genetic vulnerability and risk gene expression of Sjögren's disease for cardiovascular disease.

OBJECTIVES: This study aims to clarify the genetic associations between Sjögren's Disease (SD) and cardiovascular disease (CVD) outcomes, and to conduct an in-depth exploration of specific pleiotropic susceptibility genes. METHODS: We performed two-sample and multivariable Mendelian randomization (MR) analysis to investigate the association between SD and the risk of ischemic heart disease (IHD) and stroke. Linkage disequilibrium score regression (LDSC) and Bayesian co-localization analyses were employed to assess the genetic associations between traits. Cross-phenotype analyses were employed to identify shared variants and genes, followed by a Transcriptome-Wide Association Study (TWAS) and Multi-marker Analysis of Genomic Annotation (MAGMA) based on Multi-Trait Analysis of GWAS (MTAG) results. To validate the pleiotropic genes, we further analyzed tissue-specific differentially expressed genes (DEGs) related to SD using RNA sequencing data. RESULTS: The two-sample and multivariable MR analyses revealed that SD confers a genetic vulnerability to IHD and stroke. LDSC and co-localization analyses indicated a strong genetic linkage between SD and CVDs. Cross-phenotype analyses identified 38 and 37 pleiotropic single nucleotide polymorphisms (SNPs) for SD-Stroke and SD-IHD, respectively, primarily located within the MHC class region on 6p21.32:33 loci. Additionally, TWAS and MAGMA analyses identified pleiotropic genes located outside the MHC regions-seven associated with stroke (UHRF1BP1, SNRPC, BLK, FAM167A, ARHGAP27, C8orf12, and PLEKHM1) and two associated with IHD (UHRF1BP1 and SNRPC). Proxy variants within these genes in SD suggested an increased causal risk for stroke or IHD. Co-localization analysis further reinforced that SD and stroke share significant SNPs within the loci of FAM167A, BLK, C8orf12, SNRPC, and UHRF1BP1. DEG analysis revealed a significant up-regulation of the identified genes in SD-specific tissues. CONCLUSIONS: SD appears genetically predisposed to an increased risk of CVDs. Moreover, this research not only identified pleiotropic genes shared between SD and CVDs, but also, for the first time, detected key gene expressions that elevate CVD risk in SD patients-findings that may offer promising therapeutic targets for patient management.

Humans

Reduced R-loop abundance at proinflammatory loci: a shared epigenetic mechanism in inflammatory and metabolic diseases.

INTRODUCTION: R-loops, RNA-DNA hybrid structures with a displaced single-stranded DNA loop, are key regulators of transcriptional control, chromatin architecture, and genome stability and have emerging roles in inflammatory signaling. However, the relationship between R-loop abundance and strongly modulated inflammatory effector genes in metabolic inflammation and influenza virus infection remains underexplored. METHODS: We performed a locus-centric integrative analysis combining robust differentially expressed genes (DEGs) from multiple inflammatory and infection-related murine and human transcriptomic disease models with experimentally validated multi-cell R-loop annotations from the reference atlas RLoopBase. Our correlation framework evaluated the directional relationship between R-loop abundance and inflammatory gene expression rather than assuming disease-sample-matched R-loop measurements. We further analyzed R-loop regulatory proteins, NRF2-associated R-loop regulators, and overlaps between R-loop regulators and CRISPRi-identified mitochondrial and cellular reactive oxygen species (ROS) regulators. RESULTS: In angiotensin II-infused apolipoprotein E-deficient (ApoE-/-) mice, a model of abdominal aortic aneurysm (AAA), genomic regions encoding the top significantly upregulated genes exhibited significantly fewer R-loops than those encoding downregulated genes at days 14 and 28. Similarly, in atherosclerotic ApoE-/- mice fed a high-fat diet for 32 and 78 weeks, upregulated genes were associated with fewer R-loops than downregulated genes. Reduced R-loop abundance was also observed in genomic regions encoding the top significantly upregulated genes in liver tissues from patients with non-alcoholic steatohepatitis (NASH), as well as in monosodium urate (MSU)-stimulated lymphatic endothelial cells (LECs) and influenza virus-infected human umbilical vein endothelial cells (HUVECs). R-loop regulatory proteins upregulated during metabolic inflammation were enriched in immune and inflammatory pathways. NRF2 was identified as a regulator of 27 R-loop regulatory proteins, including 10 positively and 17 negatively regulated proteins. Furthermore, 54 R-loop regulatory proteins overlapped with CRISPRi-identified mitochondrial and cellular ROS regulators, suggesting potential reciprocal regulation between R-loop homeostasis and ROS signaling. Disease-associated changes in pro-ROS and anti-ROS R-loop regulatory proteins further linked R-loop regulation to inflammatory and oxidative stress pathways. DISCUSSION: These findings identify reduced R-loop abundance at genomic regions encoding strongly upregulated inflammatory genes as a shared feature across multiple models of metabolic inflammation and influenza virus infection. The results further suggest that immune-associated R-loop regulatory proteins and the NRF2-ROS axis may contribute to R-loop remodeling during inflammatory disease. This integrative framework provides new insight into the potential role of R-loops and ROS-sensitive R-loop regulators in inflammatory and metabolic diseases and identifies candidate pathways for future mechanistic investigation and therapeutic targeting.

R-loop regulatory proteins

Proteomic hub proteins CDKN2B, TRAPPC2L, WFS1, and ARPP19 drive biochemical recurrence and metastatic progression in prostate cancer: Protein macromolecule action.

The biological characteristics and metastasis mechanism of prostate cancer are complex, involving the important role of many proteins in cell transcriptional regulation. This study focused on the role of the proteomic hub proteins CDKN2B, TRAPPC2L, WFS1 and ARPP19 in the biochemical recurrence and metastasis progression of prostate cancer. Cross-platform transcriptome integration and differential expression analysis were used to evaluate transcriptome characteristics in a prostate cancer cohort. Functional enrichment analysis was performed by gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway annotation, and weighted gene co-expression network analysis (WGCNA) was used to investigate cancer progression subtypes. It was found that prostate cancer progression showed significant transcriptome heterogeneity, and low-expression genes dominated. We reveal the important role of epithelial-immune interactions and inflammatory signaling in transcriptional remodeling in prostate cancer. The co-expression network topology analysis showed that the immune-metabolic center module plays a central role in cancer progression. CDKN2B was identified as a key transcriptional determinant in prostate cancer typing, while TRAPPC2L and WFS1 acted as core transcriptional regulators, driving metastatic heterogeneity. ARPP19 and LOC650152 also show important transcriptional driving effects in advanced prostate cancer.

Humans

Integrative multi-omics profiling of insomnia-related molecular features reveals microbiome, immune, and therapy-relevant heterogeneity in colorectal cancer.

Emerging evidence implicates insomnia as a potential risk factor in carcinogenesis, potentially involving systemic inflammation, circadian disruption, and microbiome alterations. However, the molecular associations linking insomnia-related features to colorectal cancer (CRC), particularly with respect to tumor biology, immune microenvironmental states, and therapy-relevant phenotypes, remain largely unexplored. Multi-omics integration of genomic, transcriptomic, and microbiome data from 3,026 CRC patients across seven independent cohorts, including a large, well-annotated Clinical Omics study of Colorectal Cancer in China (COCC) cohort, enabled insomnia-based molecular classification through unsupervised non-negative matrix factorization (NMF) clustering. The insomnia subtype (IS) was biologically characterized via pathway enrichment, immune deconvolution, microbial profiling, and single-cell transcriptomics. Furthermore, an insomnia score (ISscore) was developed and validated in multiple cohorts for risk stratification and assessment of treatment-response-related indicators in CRC. Unsupervised clustering revealed two distinct molecular subtypes (IS1/IS2), with IS2 demonstrating significantly poorer survival. IS2 exhibited marked activation of EMT/angiogenesis pathways versus cell cycle activation in IS1. The IS2 microenvironment showed increased immunosuppression-related infiltration and exhausted T cell signatures, together with intratumoral microbiome variation characterized by depletion of Ruminococcaceae UCG-002 and enrichment of Hungatella/Selenomonas. The ISscore system stratified survival risk and was associated with computational indicators of immunotherapy response. Single-cell analysis nominated PPIA-BSG as a potential cell-cell communication signal involving high-ISscore tumor cells, CXCL12+ endothelial cells, and CLEC9A+ dendritic cell subsets. This multi-omics characterization of insomnia-CRC interplay suggests that insomnia-related molecular features are associated with an immunologically distinct and microbiome-altered tumor ecosystem. The ISscore provides a reproducible framework for capturing insomnia-related molecular heterogeneity, supporting risk stratification and future evaluation of therapy-relevant phenotypes.IMPORTANCEChronic insomnia affects millions, but it is not typically considered a cancer risk factor. Our study, analyzing vast biological data from over 3,000 colorectal cancer patients, uncovers a potential link between a person's predisposition to insomnia and their risk of developing this disease. This suggests that the biological pathways related to sleep may play a role in cancer development. Understanding this connection opens up new avenues for identifying individuals at higher risk and developing novel prevention strategies for colorectal cancer.

colorectal cancer

Putative function and prognostic molecular marker of mast cells in colorectal cancer.

BACKGROUND: The increased demand for markers for colorectal cancer (CRC) highlights the importance of investigating immune cells involved in CRC progression. This study aims to dissect the mast cells in CRC, characterize the role of mast cells in CRC development, coordinate molecular communication between mast cells and malignant cells, and construct and validate a prognostic classification model based on mast cell markers. METHODS: Single-cell transcriptome data of CRC patients were extracted from GSE146771 for cell classification and annotation. The malignant cells were identified by copykat and the communication between mast cells and malignant cells was analyzed by CellChat. Least absolute shrinkage and selection operator (LASSO) regression analysis and Cox regression analysis of mast cell markers were performed in the TCGA-COAD cohort to construct a prognostic classification model. qRT-PCR was performed to detect the mRNA expression of the molecules in the classification model in P815 and MC-9 cells. The co-culture experiment of MC38 and P815 cells were performed in 12-well transwell dish. Wound healing assay and Transwell assay were performed to detect cell migration and invasion. RESULTS: 10,186 high-quality cells in GSE146771 were annotated to 9 cell types. Six markers in mast cells (HDC, GATA2, ASAH1, BTBD19, TIMP1, FAM110A) were selected to construct a classification model. The high-risk score defined showed high infiltration of immunosuppressive cells, including endothelial cells, CAFs, Tregs and high angiogenesis and epithelial-mesenchymal transition (EMT) activities. In the model, HDC were abnormally low expressed in P815 cells, while BTBD19, FAM110A, GATA2, ASAH1 and TIMP1 showed excessive expression in P815 cells. Knockdown of GATA2 in the co-culture system of P815 and MC38 cells blocked cell migration and invasion. CONCLUSION: This study identified the cell types within CRC, elaborated the cellular functions of mast cells in CRC development and their molecular communication to coordinate malignant cells, and highlighted the molecular components and biological features that constitute promising prognostic classification model.

Mast Cells

Telomere-to-telomere genome of Phoebe chekiangensis reveals that age-dependent CHG hypomethylation promotes floral transition via MADS-box gene activation.

Phoebe species are renowned for their highly valuable 'golden thread' timber; however, their protracted juvenile phase presents a significant obstacle to mechanistic investigations of floral induction. Phoebe chekiangensis, a rare early-flowering representative within this genus, provides a unique model system for dissecting the vegetative-to-reproductive phase transition. Nevertheless, the absence of a high-quality reference genome has severely hindered molecular insights into its developmental regulation. Here, we present the first telomere-to-telomere (T2T) genome assembly for P. chekiangensis, comprising two completely gap-free haplotypes with contig N50 values exceeding 65 Mb, base-level quality scores >36, and Long Terminal Repeat Assembly Index scores surpassing the gold standard threshold of 20. Approximately 29 000 genes were annotated per haplotype, supported by a BUSCO completeness score of >97%. Age-resolved transcriptomic landscapes identified two MADS-box transcription factors, PcMADS5 (AP1-like) and PcMADS19.1 (SOC1-like), as core activators of the floral transition. Both genes triggered precocious flowering when ectopically expressed in Arabidopsis thaliana. Whole-genome bisulfite sequencing revealed a progressive, age-dependent decline in CHG (where H is A, C, or T) DNA methylation, which was particularly pronounced at the PcMADS19.1 locus. Notably, DML1/2, which mediate active DNA demethylation, were coordinately upregulated during the onset of reproductive growth. Chemical demethylation using 5-azacytidine further diminished CHG methylation and selectively enhanced PcMADS19.1 expression, confirming a causal relationship between CHG hypomethylation and transcriptional activation. This work delivers the first chromosome-scale T2T genome within the genus Phoebe and uncovers CHG demethylation as a previously unrecognized epigenetic switch governing reproductive competence in woody perennials.

Journal Article