Search PubMedSearch

SEARCH · Search PubMed

Results for “expression profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Identification of key immune-related genes and potential therapeutic drugs in diabetic nephropathy based on machine learning algorithms.

BACKGROUND: Diabetic nephropathy (DN) is a major contributor to chronic kidney disease. This study aims to identify immune biomarkers and potential therapeutic drugs in DN. METHODS: We analyzed two DN microarray datasets (GSE96804 and GSE30528) for differentially expressed genes (DEGs) using the Limma package, overlapping them with immune-related genes from ImmPort and InnateDB. LASSO regression, SVM-RFE, and random forest analysis identified four hub genes (EGF, PLTP, RGS2, PTGDS) as proficient predictors of DN. The model achieved an AUC of 0.995 and was validated on GSE142025. Single-cell RNA data (GSE183276) revealed increased hub gene expression in epithelial cells. CIBERSORT analysis showed differences in immune cell proportions between DN patients and controls, with the hub genes correlating positively with neutrophil infiltration. Molecular docking identified potential drugs: cysteamine, eltrombopag, and DMSO. And qPCR and western blot assays were used to confirm the expressions of the four hub genes. RESULTS: Analysis found 95 and 88 distinctively expressed immune genes in the two DN datasets, with 14 consistently differentially expressed immune-related genes. After machine learning algorithms, EGF, PLTP, RGS2, PTGDS were identified as the immune-related hub genes associated with DN. In addition, the mRNA and protein levels of them were obviously elevated in HK-2 cells treated with glucose for 24 h, as well as their mRNA expressions in kidney tissues of mice with DN. CONCLUSION: This study identified 4 hub immune-related genes (EGF, PLTP, RGS2, PTGDS), as well as their expression profiles and the correlation with immune cell infiltration in DN.

Diabetic Nephropathies

Unravelling the transcriptomic characteristics of bronchoalveolar lavage in post-covid pulmonary fibrosis.

BACKGROUND: Post-Covid Pulmonary Fibrosis (PCPF) has emerged as a significant global issue associated with a poor quality of life and significant morbidity. Currently, our understanding of the molecular pathways of PCPF is limited. Hence, in this study, we performed whole transcriptome sequencing of the RNA isolated from the bronchoalveolar lavage (BAL) samples of PCPF and compared it with idiopathic pulmonary fibrosis (IPF) and non-ILD (Interstitial Lung Disease) control to understand the gene expression profile and associated pathways. METHODS: BAL samples from PCPF (n = 3), IPF (n = 3), and non-ILD Control (n = 3) (individuals with apparent healthy lung without interstitial lung disease) groups were obtained and RNA were isolated for whole transcriptomic sequencing. Differentially Expressed Genes (DEGs) were determined followed by functional enrichment analysis and qPCR validation. RESULTS: A panel of differentially expressed genes were identified in bronchoalveolar lavage fluid cells (BALF) of PCPF as compare to control and IPF. Our analysis revealed dysregulated pathways associated with cell cycle regulation, immune responses, and neuroinflammatory processes. Real-time validation further supported these findings. The PPI network and module analysis shed light on potential biomarkers and underscore the complex interplay of molecular mechanisms in PCPF. The comparison of PCPF and IPF identified a significant downregulation of pathways that were more prominent in IPF. CONCLUSION: This investigation provides crucial insights into the molecular mechanism of PCPF and also outlines avenues for prospective research and the development of therapeutic approaches.

Humans

Identification of aquaporin (AQP) genes in the noble scallop Chlamys nobilis and characterization of their expression under low-temperature stress.

Aquaporins (AQPs) are transmembrane channel proteins essential for water homeostasis and cellular stress responses. In marine bivalves, their roles in cold tolerance remain poorly understood despite frequent winter mortality events in aquaculture. Here, we identified nine AQP genes in the genome of the economically important noble scallop Chlamys nobilis. Phylogenetic analysis revealed strong conservation with other bivalve AQPs, and structural features, including conserved NPA motifs and ar/R selectivity filters, support their canonical water/glycerol transport functions. Tissue-specific expression profiling showed predominant enrichment in osmoregulatory tissues (gills, intestine) and gonads. Under both chronic and acute low-temperature stress from 23 °C to 9 °C, most CnAQP genes exhibited transient upregulation followed by suppression. Notably, CnAQP4 displayed sustained upregulation, implicating it as a key mediator of long-term cold adaptation. Promoter analysis further revealed abundant cis-elements linked to growth and development as well as immune regulation. Our findings provide the first comprehensive characterization of the AQP family in C. nobilis, highlighting its critical role in maintaining cellular integrity during cold stress and offering molecular targets for selective breeding of cold-tolerant scallop strains.

Animals

Structural properties of short-chain carboxylic acids and alcohols relate to the molecular and physiological response of Salmonella enterica in an acidic environment.

Short-chain carboxylic acids (SCCA) and short-chain alcohols (SCALC) are naturally occurring antimicrobials that contribute to the biopreservation of food fermentations. This study investigated the effect of structurally different SCCA/SCALC with two-carbon (acetic acid; phenylacetic acid; 2-phenylethanol), three-carbon (propionic acid; 3-phenylpropionic acid; 3-phenylpropanol), and three-carbon chain with an additional hydroxyl group (lactic acid; 3-phenyllactic acid; 1-phenylpropanol) on the fitness, metabolic activity and gene expression of the pathogen Salmonella enterica at pH 4.5. SCCA inhibited Salmonella at lower concentrations than SCALC with the exception of lactic acid, which was partly consumed. The presence of a phenyl group enhanced antimicrobial activity. SCCA but not SCALC increased the lag phase of S. enterica, and in general, acetate was formed when cell growth was reduced by 20% suggesting a negative impact on bacteria fitness. Principal component analysis and hierarchical clustering indicated distinct gene expression profiles of S. enterica in response to SCCA or SCALC. In the presence of certain SCCA/SCALC, Salmonella activated pathways related to cellular pH control, and 1,2-propanediol, propionic acid and ethanolamine metabolism that involved the formation of metabolosomes. Genes related to flagellar assembly were less expressed and mobility was lower in the presence of lactic and 3-phenyllactic acid compared to controls suggesting a compound-specific response. KEY POINTS: • Differences in response among structurally different SCCA/SCALC at acidic condition. • SCCA/SCALC stress interfered with cell growth and metabolism of acetic and propionic acid. • Lactic acid prolonged the lag phase and reduced motility of Salmonella.

Salmonella enterica

Identification and analysis of HD-ZIP transcription factors that regulate salt gland development and salt tolerance in Limonium bicolor.

Soil salinity severely constrains agricultural production. Elucidating the salt-tolerance mechanisms of halophytes can provide innovative approaches for improving the salt tolerance of crop plants. In this study, we performed genome-wide identification and analysis of 36 LbHDZ genes encoding homeodomain-leucine zipper (HD-ZIP) transcription factors in Limonium bicolor, a typical recretohalophyte that excretes excess salt ions through specialized salt glands. Expression profiling across different stages of salt gland development, as well as in various tissues under salt stress, indicated that multiple LbHDZ genes are involved in regulating salt gland development and salt tolerance. Among these genes, LbHDZ14 (a member of the HD-ZIP II subfamily) exhibited sustained high expression during the critical period of salt gland formation, while its transcript levels were significantly downregulated in leaves and roots under salt stress. Subsequent experiments demonstrated that LbHDZ14 is localized in the nucleus and negatively regulates salt gland density and salt tolerance by directly binding to the promoter of LbGDSL, a positive regulator of salt gland development. In conclusion, this study reveals the expression patterns of LbHDZ genes in L. bicolor, characterizes the functional mechanism of LbHDZ14, further elucidates the regulatory network underlying salt gland development, and provides candidate genes for enhancing crop salt tolerance.

Plumbaginaceae

A nucleolar stress gene signature enables quantitative scoring across multi-omics contexts.

The nucleolus is essential for ribosome biogenesis and cellular homeostasis, and its dysfunction can induce nucleolar stress, a process implicated in cancer and other diseases. However, nucleolar stress is commonly inferred from morphological changes or a limited set of functional assays, and quantitative approaches based on gene expression profiles remain lacking. Here, we integrate literature curation with multi-dataset screening to define a nucleolar stress gene signature and develop a nucleolar stress score (NuS) applicable to bulk transcriptomics, single-cell transcriptomics, proteomics, and spatial transcriptomics. Using this framework, we show in colorectal cancer models that oxaliplatin induces nucleolar stress, suppresses nascent rRNA synthesis, and activates p53 signaling, whereas these responses are attenuated in oxaliplatin-resistant cells. Combined with a ribosome biogenesis activity score (RiboSis), NuS captures related but distinct dimensions of nucleolar function and stratifies tumors into functional states associated with clinical outcomes. NuS-based analysis of perturbational transcriptomes further prioritizes compounds with putative nucleolar stress-inducing activity. Collectively, this study provides a quantitative framework for evaluating nucleolar stress and illustrates its applications in disease stratification and drug mechanism discovery.

Cell Nucleolus

Integrative TWAS and multi-omics analyses prioritize HSPE1 as a candidate risk gene for bipolar disorder with immune cell-specific regulatory evidence.

BACKGROUND: Bipolar disorder (BD) is a severe psychiatric disorder associated with substantial disability. Although genome-wide association studies have identified multiple BD-associated loci, the underlying genes and mechanisms remain incompletely understood. METHODS: We integrated a European-ancestry BD genome-wide association dataset with cross-tissue and tissue-specific transcriptome-wide association studies (TWAS) and complementary gene-based analysis. Candidate genes were further evaluated using differential expression analysis, consensus clustering, immune infiltration analysis, machine learning, summary-data-based Mendelian randomization, Mendelian randomization using single-cell expression quantitative trait locus data, single-nucleus transcriptomics, phenome-wide association analysis, and virtual screening. RESULTS: The integrative analyses prioritized 37 candidate genes. Peripheral-blood differential-expression analysis identified 14 genes that remained significant after FDR correction, and their expression profiles separated BD samples into two expression-defined clusters. Machine-learning analysis selected UNC50, LMAN2L, LYG2, HSPE1, and KANSL3 for an exploratory classification nomogram. SMR associated genetically predicted higher HSPE1 expression with increased BD risk in two blood eQTL datasets. Cell-type-specific analyses indicated HSPE1-related associations in T-cell and natural killer cell subsets, while single-nucleus analysis descriptively showed higher HSPE1 expression in medial thalamic T cells from BD samples. PheWAS identified no genome-wide significant associations for HSPE1, whereas virtual screening identified candidate compounds with favorable predicted docking scores against the HSPE1 structure. CONCLUSION: This integrative multi-omics study identified HSPE1 as a candidate BD risk gene with immune-cell-related regulatory evidence, providing insight into BD pathogenesis and supporting functional validation.

Humans

Chicken vitellogenin gene-binding protein, a leucine zipper transcription factor that binds to an important control element in the chicken vitellogenin II promoter, is related to rat DBP.

We screened a chicken liver cDNA expression library with a probe spanning the distal region of the chicken vitellogenin II (VTGII) gene promoter and isolated clones for a transcription factor that we have named VBP (for vitellogenin gene-binding protein). VBP binds to one of the most important positive elements in the VTGII promoter and appears to play a pivotal role in the estrogen-dependent regulation of this gene. The protein sequence of VBP was deduced from a nearly full length cDNA copy and was found to contain a basic/zipper (bZIP) motif. As expected for a bZIP factor, VBP binds to its target DNA site as a dimer. Moreover, VBP is a stable dimer free in solution. A data base search revealed that VBP is related to rat DBP. However, despite the fact that the basic/hinge regions of VBP and DBP differ at only three amino acid positions, the DBP binding site in the rat albumin promoter is a relatively poor binding site for VBP. Thus, the optimal binding sites for VBP and DBP may be distinct. Similarities between the VBP and DBP leucine zippers are largely confined to only four of the seven helical spokes. Nevertheless, these leucine zippers are functionally compatible and appear to define a novel subfamily. In contrast to the bZIP regions, other portions of VBP and DBP are markedly different, as are the expression profiles for these two genes. In particular, expression of the VBP gene commences early in liver ontogeny and is not subject to circadian control.

Amino Acid Sequence

GeneCOCOA: Detecting context-specific functions of individual genes using co-expression data.

Extraction of meaningful biological insight from gene expression profiling often focuses on the identification of statistically enriched terms or pathways. These methods typically use gene sets as input data, and subsequently return overrepresented terms along with associated statistics describing their enrichment. This approach does not cater to analyses focused on a single gene-of-interest, particularly when the gene lacks prior functional characterization. To address this, we formulated GeneCOCOA, a method which utilizes context-specific gene co-expression and curated functional gene sets, but focuses on a user-supplied gene-of-interest (GOI). The co-expression between the GOI and subsets of genes from functional groups (e.g. pathways, GO terms) is derived using linear regression, and resulting root-mean-square error values are compared against background values obtained from randomly selected genes. The resulting p values provide a statistical ranking of functional gene sets from any collection, along with their associated terms, based on their co-expression with the gene of interest in a manner specific to the context and experiment. GeneCOCOA thereby provides biological insight into both gene function, and putative regulatory mechanisms by which the expression of the GOI is controlled. Despite its relative simplicity, GeneCOCOA outperforms similar methods in the accurate recall of known gene-disease associations. We furthermore include a differential GeneCOCOA mode, thus presenting the first implementation of a gene-focused approach to experiment-specific gene set enrichment analysis. GeneCOCOA is formulated as an R package for ease-of-use, available at https://github.com/si-ze/geneCOCOA.

Gene Expression Profiling

Genome-Wide Identification of the PAL Gene Family in Idesia polycarpa and Transcriptomic Responses to Botryosphaeria dothidea Infection.

Idesia polycarpa is a woody oil tree threatened by stem canker caused by Botryosphaeria dothidea, yet the organization and infection-responsive behavior of its phenylalanine ammonia-lyase (PAL) gene family remain poorly understood. Here, we identified five IpPAL genes and characterized their phylogenetic relationships, conserved features, duplication patterns, promoter cis-elements, and infection-associated expression profiles. Segmental and tandem duplication contributed to IpPAL family evolution, and all duplicated pairs showed Ka/Ks ratios below 1, consistent with purifying selection. RNA sequencing (RNA-seq) of contrasting Chengdu and Zhangjiajie provenances revealed distinct temporal responses. In Chengdu, IpPAL2-IpPAL4 were significantly upregulated at 24 h after inoculation, whereas all five genes were upregulated at 96 h. In Zhangjiajie, all five IpPAL genes were significantly upregulated at 24 h, while IpPAL2-IpPAL5 remained upregulated at 96 h. No IpPAL gene met the differential-expression criteria between provenances under mock conditions or at 24 h; at 96 h, IpPAL1 and IpPAL3 were lower and IpPAL5 was higher in Zhangjiajie than in Chengdu. Scanning electron microscopy (SEM) provided complementary qualitative evidence of provenance-associated tissue responses. These findings demonstrate time- and gene-specific IpPAL responses to B. dothidea and identify candidate genes for further functional analysis.

Ascomycota

Genome-wide DNA methylation and transcriptome sequencing analyses of lens tissue in an age-related mouse cataract model.

DNA methylation is known to be associated with cataracts. In this study, we used a mouse model and performed DNA methylation and transcriptome sequencing analyses to find epigenetic indicators for age-related cataracts (ARC). Anterior lens capsule membrane tissues from young and aged mice were analyzed by MethylRAD-seq to detect the genome-wide methylation of extracted DNA. The young and aged mice had 76,524 and 15,608 differentially methylated CCGG and CCWGG sites, respectively. The Pearson correlation analysis detected 109 and 33 differentially expressed genes (DEGs) with negative methylation at CCGG and CCWGG sites, respectively, in their promoter regions. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analyses showed that DEGs with abnormal methylation at CCGG sites were primarily associated with protein kinase C signaling (Akap12, Capzb), protein threonine kinase activity (Dmpk, Mapkapk3), and calcium signaling pathway (Slc25a4, Cacna1f), whereas DEGs with abnormal methylation at CCWGG sites were associated with ribosomal protein S6 kinase activity (Rps6ka3). These genes were validated by pyrosequencing methylation analysis. The results showed that the ARC group (aged mice) had lower Dmpk and Slc25a4 methylation levels and a higher Rps6ka3 methylation than the control group (young mice), which is consistent with the results of the joint analysis of differentially methylated and differentially expressed genes. In conclusion, we confirmed the genome-wide DNA methylation pattern and gene expression profile of ARC based on the mouse cataract model with aged mice. The identified methylation molecular markers have great potential for application in the future diagnosis and treatment of ARC.

Animals

Discovery of novel diagnostic biomarkers of hepatocellular carcinoma associated with immune infiltration.

OBJECTIVE: Diagnosis of hepatocellular carcinoma (HCC) remains challenging for clinicians. Machine learning approaches and big data analyses are viable strategies for identifying HCC diagnostic markers. MATERIALS AND METHODS: In this study, we downloaded mRNA expression profiles of HCC from the GEO database and used random forest and machine learning algorithms, such as least absolute shrinkage and selection operator, to screen for reliable diagnostic genes. Disease Ontology, Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Set Enrichment Analysis enrichment analyses were performed to explore differential gene functions and disease pathways. CIBERSORT was performed to calculate the immune cell infiltration of HCC and the correlation between diagnostic genes and immune cells. Cell experiments were performed to evaluate the function of R-spondin 3 (RSPO3) in HCC cells. Immunohistochemical staining was used to evaluate the protein expression of CD138, CD206 and iNOS. RESULTS: The results indicated that extracellular matrix protein 1 (ECM1), Niemann-Pick C1-Like 1 (NPC1L1) and RSPO3 were down-regulated in HCC compared with the normal group (p&#x2009;<&#x2009;0.05), which was validated in clinical tissue samples. Moreover, ECM1, NPC1L1 and RSPO3 had high diagnostic values (AUC > 0.75) for HCC in both training and test groups. Immuno-infiltration analysis revealed that ECM1 and RSPO3 were highly positively correlated with neutrophil and macrophage M2 levels, whereas they were negatively correlated with Tregs. RSPO3-si affected cell proliferation and apoptosis in HCC. Furthermore, RSPO3 exhibited a positive correlation with tumour progression, the proportion of plasma cells and M2 macrophages in mice, while showing a negative association with M1 macrophages. CONCLUSION: The present study identified ECM1, NPC1L1 and RSPO3 as new diagnostic biomarkers for HCC based on normal and diseased samples from HCC, meanwhile the pro-oncogenic function of RSPO3 and its regulation on immune infiltration have been confirmed.

Carcinoma, Hepatocellular

The ASH HematOmics Program supports integrative analysis of genomic and clinical data in hematologic diseases.

The increasing availability of genomic and transcriptomic sequencing has uncovered diverse genomic alterations and distinct gene expression profiles driving hematologic diseases, yet a data integration and sharing platform dedicated to hematology remains lacking. We developed the American Society of Hematology (ASH) HematOmics Program (ASHOP; ashop.hematology.org), a resource for exploring somatic alterations and gene fusions, transcriptomic results, and clinical data from 5960 patients spanning B-cell precursor and T-cell acute lymphoblastic leukemia, acute myeloid leukemia, myelodysplastic syndromes, and chronic lymphocytic leukemia. Users can explore genomic alteration landscapes and comutation patterns via lollipop and matrix plots and analyze significantly altered genes in user-defined subcohorts. Transcriptomes can be explored through interactive uniform manifold approximation and projections, clustering, differential expression, and pathway enrichment. Genomic, transcriptomic features, and clinical outcomes can be correlated in a user-driven manner or combined to precisely define study cohorts. We illustrate the following 4 use cases of ASHOP: (1) stratification of DUX4-rearranged B-cell leukemias into Early/Multipotent and Committed subgroups with distinct outcomes, (2) characterization of HOXA/HOXB expression patterns in acute myeloid leukemias, (3) correlating mutational burden with mismatch repair deficiency and mutational signatures, and (4) investigation of TP53 alteration landscape. ASHOP is an open-access resource to inform genomic and transcriptomic data interpretation for hematologic malignancies and will expand to support additional diseases and data modalities from the ASH community.

Humans

CeLLTra: aligning cell names with gene expression via a pathway-informed transformer.

MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans

Advances in the diagnosis and classification of B-ALL: comparative insights from updated guidelines.

Accurate molecular classification is essential for diagnosis, risk stratification, and treatment selection in B-cell lymphoblastic leukemia (B-ALL). In this study, we performed a comprehensive, real-world reclassification of 1015 consecutively diagnosed B-ALL patients using the fifth edition of the World Health Organization Classification of Haematolymphoid Tumours (WHO-HAEM5) and the International Consensus Classification (ICC). An integrative genomic strategy that combined whole transcriptome sequencing, fusion detection, mutational analysis, and cytogenetics enabled reclassification according to both the WHO-HAEM5 and ICC frameworks, thereby substantially reducing the proportion of unclassifiable B-ALL from 41.9% (2016 WHO revision [WHO-HAEM4R]) to 15.9% (WHO-HAEM5) and 11.9% (ICC). Distinct clinical and prognostic features were identified across newly defined subtypes. Multivariable analysis confirmed that this genomic classification is a robust, independent predictor of survival after adjusting for age, minimal residual disease status, and transplant intervention. Specifically, HLF-rearranged and MEF2D-rearranged B-ALL conferred a persistently poor prognosis across all age groups despite allogeneic hematopoietic stem cell transplantation, highlighting an urgent need for novel therapeutic strategies. Gene expression profiling resolved cryptic subtypes, including ETV6::RUNX1-like, ZNF384-rearranged-like, and BCR::ABL1-like B-ALL, and uncovered diagnostic ambiguity in patients with concurrent lesions. In addition, we report emerging high-risk groups, including IDH1/2- and ZEB2 Q1072-mutated B-ALL, that may warrant recognition as distinct molecular entities. Our findings demonstrate the clinical use of integrative transcriptomic profiling in refining B-ALL taxonomy in guiding risk-adapted therapies and informing future revisions of diagnostic standards. This study supports the incorporation of high-throughput molecular diagnostics into routine leukemia classification and precision treatment planning.

Humans