Search PubMedSearch

SEARCH · Search PubMed

Results for “dimensional reduction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Mechanisms of enhanced or impaired DNA target selectivity driven by protein dimerization.

Successful DNA transcription demands coordination between proteins that bind DNA while simultaneously binding to one another to form dimers or higher-order complexes. For proteins with numerous DNA targets throughout the genome, measurements that report on their dwell time or occupancy thus represent a convolution over a population interacting with specific DNA, nonspecific DNA, or protein partners on DNA. Dimerization is known to add contacts that can help a single protein to stably bind DNA. However, we show here that dimerization can also impair measured dwell times and occupancy on target sequences because the population redistributes across DNA. We combine mass-action kinetic models of pairwise reversible reactions between proteins and DNA with theory and spatial stochastic simulations to isolate the role of dimerization on observed DNA dwell times, occupancy, and spatial distribution of proteins on DNA. Three key themes emerge: (i) Protein-protein interactions, in addition to protein-DNA interactions, can localize a protein to DNA, and relative binding rates can thus widely tune dwell times. (ii) Dimensional reduction achieved through nonspecific binding and subsequent 1D diffusion controls the order-of-magnitude of enhancements despite nucleosome barriers. (iii) Dimerization enhances selectivity for locally clustered targets and often impairs binding to widely-spaced targets by sequestration. Compared with ChIP-seq data, our model explains how the distribution of the essential GAF protein throughout the genome is highly selective for clustered targets due to protein interactions. This model framework predicts when even weak dimerization can redistribute and stabilize proteins on DNA as a necessary part of transcription.

DNA binding

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Methylome Profiling of Cartilage Tumors: A Promising New Diagnostic Tool?

DNA methylation and copy number variation (CNV) profiling has emerged as a promising tool for the classification of bone and soft tissue tumors. We evaluated its utility in cartilage tumors, where distinguishing low-grade from high-grade conventional central chondrosarcomas (CSs) and atypical cartilaginous tumors (ACTs) from enchondromas (ECs) is a frequent diagnostic challenge, particularly on biopsy material. We analyzed 214 chondrogenic tumors, including ECs, ACTs, conventional CSs, dedifferentiated chondrosarcomas (DDCSs), and clear cell CSs, and determined their IDH1/2 mutation status. Unsupervised dimensionality reduction of genome-wide DNA methylation patterns revealed 4 clusters among IDH-mutant (MUT) tumors (IDH-MUT-1: mostly ECs and ACTs and some high-grade CSs; IDH-MUT-2: predominantly high-grade CSs; IDH-MUT-3: largely DDCSs; and IDH-MUT-SB: distinct skull base group with a markedly different methylation pattern) and 2 clusters among IDH-wild-type (WT) tumors (IDH-WT-1 and IDH-WT-2: both primarily high-grade CSs, with IDH-WT-2 showing higher tumor grade and more extensive CNVs). Clear cell CSs formed a separate cluster. The amount of CNVs, including loss of CDKN2A, increased with tumor grade, reflecting increased genomic instability during chondrosarcoma progression. Supervised classifiers trained separately, both on methylation and CNV data, and distinguished low-grade and high-grade cartilaginous tumors with area under the curve values of 0.87 to 0.97 and 85% to 90% accuracy. Furthermore, we tested whether DDCSs can be distinguished from metastatic carcinomas and other high-grade sarcomas of the bone. Across 246 reference samples, a supervised classifier achieved 97.2% accuracy (area under the curve, 99.8%) and correctly identified 30 of 32 DDCSs (93.8%). These results indicate that DNA methylation and CNV data analysis provide a valuable tool for distinguishing most low- and high-grade CSs, with additional utility also in differentiating DDCS from morphologic mimics.

cartilaginous tumors

Genetic heterogeneity affects the risk of incident depression, comorbidity, and response to environment: A prospective trajectory study.

BACKGROUND: Depression exhibits significant heterogeneity in its genetic underpinnings. The role of genetic components in the development of depression and its comorbidities remains insufficiently explored. METHODS: First, depression risk loci from a large-scale genome-wide meta-analysis were annotated to Gene Ontology (GO) terms by functional enrichment. GO-based polygenic risk scores (GO-PRS) were then calculated for individuals in the UK Biobank. Principal component analysis (PCA) was applied for dimensionality reduction, followed by cluster analysis to identify genetic subtypes of depression. Multistate models were applied to assess the impact of genetic patterns on the trajectory from healthy status to incident depression, and depression to 26 subsequent diseases, as well as the associations between environmental factors and disease trajectories across genetic subtypes. RESULTS: Participants were categorized into three genetic subtypes: immune-dominant, neuro-dominant, and comprehensive-risk. Significant differences in risk of depression and subsequent diseases, and susceptibility to environmental factors were observed across subtypes. Comprehensive-risk subtype showed higher risks of depression compared to immune-dominant (HR: 1.10, 95% CI: 1.05-1.15) and neuro-dominant subtype (HR: 1.12, 95% CI: 1.08-1.16). Comprehensive-risk subtype exhibited higher risks of transition from depression to subsequent diseases, such as anemia compared to immune-dominant subtype, and diseases of the digestive system compared to neuro-dominant subtype. Environmental factors were more strongly associated with the transition from depression to subsequent diseases in immune-dominant and comprehensive-risk subtypes, including cardiovascular, respiratory, and metabolic diseases. CONCLUSIONS: Our findings highlight the genetic heterogeneity of depression and comorbidities, and shed light on how genetic components modulate responses to environmental factors.

Humans

Single-Cell Proteomics Reveals Proteome Remodeling and Cellular Heterogeneity During NGF-Induced PC12 Neuronal Differentiation.

Single-cell proteomics enables direct measurement of cellular heterogeneity during dynamic biological processes, but its application to fragile and highly adherent neuronal models remains challenging. Here, we developed and applied an optimized single-cell proteomics workflow to characterize proteome remodeling during nerve growth factor (NGF)-induced differentiation of PC12 cells. To enable reliable single-cell analysis, we implemented gentle dissociation, antiaggregation strategies, and thermal inkjet-based cell dispensing, achieving high accuracy in single-cell isolation. Inclusion of n-dodecyl-β-d-maltoside (DDM) improved recovery of membrane-associated and low-solubility proteins. Coupled with LC-ion mobility-mass spectrometry, this workflow enabled quantification of 2,000-3,000 proteins per cell across the differentiation time course. Single-cell proteomic analysis revealed progressive and heterogeneous proteome remodeling during differentiation. While undifferentiated cells formed a relatively homogeneous population, later stages (Days 4-6) exhibited increased variability, including multimodal protein abundance distributions and separation into distinct subpopulations. Dimensionality reduction, clustering, and non-negative matrix factorization identified multiple coexisting proteomic states within the same time points, reflecting asynchronous differentiation trajectories. These subpopulations were characterized by coordinated differences in pathways related to intracellular trafficking, protein translation, cytoskeletal organization, and neuronal maturation. Comparison with bulk proteomics demonstrated that proteins associated with differentiated neuronal states, including those involved in neurite formation and structural remodeling, are underrepresented in population-averaged measurements but are enriched within specific single-cell subpopulations. Temporal and cluster-resolved analyses further revealed distinct protein expression trajectories, including early decreases in cell cycle and metabolic pathways and later increases in neuronal structural and regulatory proteins. Together, this study establishes an optimized workflow for single-cell proteomics of neuronal systems and demonstrates that NGF-induced PC12 differentiation proceeds through heterogeneous and divergent proteomic states that are not resolved by bulk analysis.

Animals

Association of AGER genetic variants with chronic obstructive pulmonary disease susceptibility in Southern Chinese Han populations.

OBJECTIVE: Chronic obstructive pulmonary disease (COPD) remains a leading cause of disability and mortality among elderly populations. Studies indicate that AGER plays a critical regulatory role in the pathogenesis of respiratory disorders. However, the genetic variations in AGER to COPD susceptibility remain incompletely understood. This study employs a case-control design to investigate associations between AGER genetic variants and COPD risk in the Southern Chinese Han population. METHODS: This study enrolled 270 COPD patients and 271 healthy controls. AGER single-nucleotide polymorphisms (SNPs) were analysed using the MassARRAY iPLEX platform. Logistic regression models evaluated associations between AGER polymorphisms and COPD susceptibility, with false discovery rate (FDR) correction applied to mitigate multiple testing errors. SNP-SNP interactions were investigated through multifactor dimensionality reduction (MDR) analysis. Expression quantitative trait locus (eQTL) data from the GTEx database were further analysed to assess regulatory relationships between SNPs and AGER gene expression levels. RESULTS: This study showed that rs3134941 (G allele, OR = 0.21, 95% CI = 0.10-0.41, p (FDR) = 0.001) and rs3131300 (G allele, OR = 0.32, 95% CI = 0.20-0.49, p (FDR) = 0.0001) were significantly associated with a reduced susceptibility to COPD. MDR indicated that rs3131300 was the optimal predictive model for COPD risk. Additionally, initial mechanistic investigations utilizing the GTEx database identify rs3134941 (C > G) and rs3131300 (A > G) as significant expression quantitative trait loci for AGER mRNA in cell-cultured fibroblasts and whole blood. CONCLUSION: Our study demonstrated that AGER genetic variants might play a protective role in the progression of COPD.

Aged

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

Genomic hallmarks of depot medroxyprogesterone acetate-associated meningiomas.

BACKGROUND: Population-based studies have linked progestin exposure to increased meningioma risk. However, the molecular basis of meningiomas associated with depot medroxyprogesterone acetate (DMPA)-a common injectable contraceptive-remains undefined. METHODS: We performed an integrated clinicopathologic and genomic analysis of meningiomas from 10 women with long-term DMPA exposure. Tumors underwent histopathological analysis, targeted sequencing, and DNA methylation profiling. Data were integrated with reference cohorts (Baylor and Heidelberg) and analyzed through classifier assignment, consensus clustering, copy number analysis, differential methylation testing, and dimensionality reduction. RESULTS: Depot medroxyprogesterone acetate-associated meningiomas were all newly diagnosed, World Health Organization grade 1 tumors with a predilection for the anterior and central skull base (n = 6). Nine patients harbored multiple meningiomas. Four experienced regression of untreated meningiomas following DMPA cessation, while 5 demonstrated stabilization. Histopathology demonstrated relative overrepresentation of metaplastic morphology, an uncommon meningioma subtype. All DMPA-associated meningiomas mapped to benign molecular groups, and most exhibited low copy number alteration burden. Targeted sequencing revealed enrichment for TRAF7 mutations (n = 5), with no NF2 mutations detected. Eight tumors shared consensus cluster identity, with cohesive grouping on principal component analysis and t-distributed stochastic neighbor embedding. No differential methylation was identified at the progesterone receptor locus. CONCLUSIONS: Depot medroxyprogesterone acetate-associated meningiomas represent a recognizable phenotype within the broader NF2-wildtype/TRAF7-enriched spectrum of benign meningiomas, characterized by chromosomal stability, a shared methylation profile, tumor multiplicity, and regression or stabilization following DMPA cessation. While derived from a small single-institution cohort, these findings provide a molecular framework for understanding progestin-associated meningioma biology, reinterpreting epidemiologic literature, and informing population-level risk stratification.

Humans

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm⁻¹ wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis

Interaction between toll-like receptor 4 polymorphism and abdominal obesity on ovarian cancer risk in Chinese women.

OBJECTIVES: the aim of this study was to evaluate the impact of TLR4 gene single nucleotide polymorphisms (SNPs) and additional TLR4 gene SNP- SNP and SNP- abdominal obesity (AO) interaction on ovarian cancer (OC) risk. METHODS: Generalized multifactor dimensionality reduction method were utilized to identify the most informative interactions between four SNPs in the TLR4 gene and abdominal obesity. Logistic regression was employed to investigate the association between 4 SNPs within TLR4 gene and OC risk, and additional SNP- SNP and gene- AO interaction on OC risk, ORs (95%CI) were calculated. RESULTS: The analysis of logistic regression indicated a markedly elevated risk of OC in individuals carrying either the rs4986790-G or rs11536889-C alleles in the TLR4 gene compared to those with the standard genetic variations, adjusted ORs (95%CI) were 1.61 (1.28-1.96) and 1.48 (1.09-1.91). GMDR analysis indicated a significant two-locus model (p = 0.018) involving rs4986790 and rs11536889, and a significant two-locus model (p = 0.001) involving rs4986790 and AO. Participants with rs4986790- AG/GG and rs11536889GC/ CC genotype has the highest OC risk, compared to participants with rs4986790-AA and rs11536889-GG genotype, OR (95%CI) = 2.58 (1.46-3.71), and abdominal obese participants with rs4986790- AG/GG genotype have the highest OC risk, compared to non- abdominal obese participants with rs4986790-AA genotype, OR (95%CI) = 3.17 (1.78-4.58). CONCLUSIONS: The findings suggested that TLR4 gene rs4986790 and rs11536889 polymorphisms were associated with increased OC risk. Significant interaction also existed between rs4986790 and AO, which means that the WC levels may influence the impact of rs4986790 on OC risk.

Adult

ABO exon polymorphisms are related to ischemic stroke in a Chinese Han population.

BACKGROUND: Recent research have underscored the relation of ABO blood group system to cerebrovascular disorders predisposition. The present investigation endeavors to delve into the relationship between ABO polymorphisms and ischemic stroke (IS) risk. METHODS: A cohort of 646 IS patients and 649 matched healthy controls was recruited. Genotyping of five SNPs within ABO were conducted by Agena MassARRAY platform. Logistic regression models were employed to estimate odds ratios (ORs) and 95% confidence intervals (CIs). Additionally, SNP-SNP interaction was assessed by multifactor dimensionality reduction (MDR) method. Furthermore, Analysis of Variance (ANOVA) was utilized to explore the association between genotypes and blood lipid profiles. RESULTS: The study identified an elevated IS risk associated with rs8176740 and rs8176720 in the overall population. Notably, ABO rs8176720 emerged as the most informative single-locus model for IS susceptibility. These variants were related to an elevated IS risk, specifically in female subjects, the subgroup aged > 64 years, non-smokers, drinkers or non-drinkers. Moreover, rs8176749 and rs8176745 were associated with red blood cell count levels and total bilirubin levels. CONCLUSION: This study firstly demonstrated the association of ABO rs8176740 and rs8176720 with IS incidence, which increased the understanding regarding the effect of ABO on IS pathogenesis.

Aged

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis

Myeloid landscape of BRAF-mutant papillary thyroid cancer and thyroiditis.

Papillary thyroid cancer (PTC) is less aggressive when associated with lymphocytic thyroiditis (LT), even in the presence of oncogenic BRAF, including smaller tumours, less lymph node involvement and reduced extrathyroidal extension. To investigate possible immune mechanisms underlying this association, we compared the tumour microenvironment of PTC-BRAF with LT and that without LT using single-cell RNA sequencing (scRNA-seq). Single-cell libraries were generated from fresh and fixed tumour samples with post-dissociation viability >70% using the 10x Genomics Chromium Platform and sequenced on an Illumina NovaSeq 6000. We analysed scRNA-seq data from 11 PTC-BRAF tumours: four with LT (one publicly available sample) and seven without LT. Downstream analyses included quality control, batch correction, dimensionality reduction, and differential gene expression analysis. We found that neutrophils were the predominant myeloid cell type in PTCs without LT. Thyrocytes without LT showed significant expression of the neutrophil recruitment chemokine ECRG4. In the absence of LT, neutrophils expressed oncogenic genes with poor clinical outcomes. In contrast, thyrocytes from tumours with LT showed increased expression of MHC-II antigen presentation, consistent with effective immune surveillance. Thyrocytes and macrophages in the presence of LT showed enrichment of interferon gamma response pathways. Our data suggest that LT in thyroid cancer is associated with enhanced antigen presentation and fewer features of pro-tumourigenic innate immune activity. These results identify previously under-recognised innate immune cell population and associated transcriptomic features, which suggest new mechanisms to target immune treatments in PTC refractory to other therapies.

Humans

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

Systematic background selection with BasCoD enhances contrastive dimension reduction in single cell genomics.

In single-cell experiments spanning diverse conditions, distinguishing variation specific to one condition (e.g., treatment) from shared or background variation (e.g., control) is critical for uncovering treatment-specific molecular responses. However, these studies typically yield ultra-high-dimensional data, necessitating effective dimension reduction for reliable biological interpretation. Contrastive dimension reduction methods address this challenge by identifying low-dimensional features enriched in a target dataset relative to a background dataset that captures shared variation. Despite their growing utility, the success of such methods critically depends on the choice of background, yet no formal criterion exists for evaluating or selecting backgrounds. To address this gap, we introduce BasCoD, a statistical testing framework based on spectral subspace inclusion theory, that enables rigorous evaluation and systematic selection of background datasets. Applying BasCoD across a range of single-cell datasets, we show that it effectively identifies suitable backgrounds, substantially improving the contrast and interpretability of the resulting target representations. We further demonstrate how BasCoD can guide the design of contrastive analyses in large-scale single-cell experiments conducted under heterogeneous conditions and elucidate potential interaction effects in perturbation studies.

Single-Cell Analysis