Search PubMedSearch

SEARCH · Search PubMed

Results for “differential analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem

A Comprehensive Analysis of Differential Protein Expression in the Plasma of Rheumatoid Arthritis Patients Utilizing Data-Independent Acquisition (DIA) Proteomics Technology.

BACKGROUND: Rheumatoid Arthritis (RA) is a Prevalent Autoimmune Disorder Affecting Millions of People Worldwide. A Thorough Understanding of Its Clinical and Pathological Features Is Essential to Improve Patient Outcomes. METHODS: This Study Combined Data-Independent Acquisition Proteomics and Enzyme-Linked Immunosorbent Assay (ELISA) to Identify and Validate Potential Plasma Protein Biomarkers for the Early Diagnosis of RA. RESULTS: Differential Proteomic Analysis Identified Differentially Expressed Proteins Between Patients With RA and Healthy Controls and Characterized Their Functions. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes Enrichment Analyses Were Performed to Explore Protein Functions and Associated Biological Pathways. The STRING Database and the Metascape Platform Were Used to Conduct an in-Depth Analysis of the Protein-Protein Interaction Network, Highlighting the Functional Attributes and Interconnections of Upregulated Proteins and Identifying Key Protein Complexes Involved in RA. ELISA Analysis of Plasma Samples Revealed Significantly Elevated SERPINA3 Levels in Patients With RA, Which Were Positively Correlated With Disease Activity Indicators-Including Erythrocyte Sedimentation Rate, C-Reactive Protein, and Disease Activity Score 28-But Were Not Correlated With Rheumatoid Factor or Its Subtypes. CONCLUSIONS: This Study Provides New Insights and Identifies Potential Biomarkers for the Early Diagnosis of RA.

Humans

Identification of ultrasound-associated gene candidates in myeloid cells and construction of a prognostic risk model for acute myeloid leukemia.

BACKGROUND: Incorporating ultrasound (US) treatment sensitivity analysis may improve the treatment of acute myeloid leukemia (AML). METHODS: This study integrated single-cell and bulk datasets for analysis. Differential expression analysis between US-treated and control samples was performed using limma package. The AUCell package was used to calculate US-associated scores in the single-cell dataset. Differentially expressed genes (DEGs) between the specific groups were identified, followed by intersection analysis with previously identified DEGs. Univariate regression, Least Absolute Shrinkage and Selection Operator (LASSO) analysis (using the glmnet package), and stepwise multivariate regression (using the MASS package) were used to refine the candidate genes and to construct a risk model. The model genes were validated using in vitro experiments. Enrichment analysis was conducted using gene set enrichment analysis (GSEA), and immune infiltration was evaluate by single-sample GSEA (ssGSEA) and ESTIMATE algorithms. The correlations between RiskScores and drug sensitivity were analyzed by oncoPredict package. Finally, tumor mutational burden (TMB) and genomic mutations were compared between the risk groups. RESULTS: Nine prognostic signatures (SPINK2, HNRNPAB, SH3BGRL3, CLEC11A, ITGA4, RPL39L, MX1, HEXIM1, and MAP4K4) were identified. Particularly, low expression of SPINK2 attenuated the activity and invasion of AML cells. High-risk group had higher immune cell infiltration. Eight drugs were predicted to be correlated with the RiskScore model. DNMT3A and RUNX1 showed higher mutation frequencies in the high-risk group, whereas KIT and MUC16 showed higher mutation frequencies in the low-risk group. CONCLUSION: The RiskScore model established in this study provides a theoretical basis for clinically screening responsive populations and optimizing treatment strategies.

Humans

Transcriptomic responses to developmental temperature in two field-collected Spodoptera exigua populations from Korea.

The beet armyworm, Spodoptera exigua, is a polyphagous insect whose development and seasonal occurrence are strongly influenced by temperature. However, transcriptomic responses to developmental thermal regimes remain insufficiently characterized in field-collected populations. In this study, we compared two Korean field-collected populations of S. exigua: a Haenam population collected in May and initially maintained at 15 ± 1 °C (HN), and a Jeju population collected in July and initially maintained at 27 ± 1 °C (JJ). F1 larvae from each population were reared under three fluctuating developmental temperature regimes: low (15-21 °C), middle (21-27 °C), and high (27-33 °C), followed by RNA-seq analysis. Differential expression analysis revealed population-associated variation in transcriptomic responses across developmental temperatures. HN exhibited a larger number of differentially expressed genes under the high-temperature regime, suggesting stronger transcriptomic sensitivity to elevated developmental temperature. Functional enrichment analyses identified population-associated differences in pathways related to heat response, oxidative metabolism, cytoskeletal organization, cuticle-associated processes, lipid metabolism, and immune-related functions. In JJ, heat-response and cuticle-related expression patterns were more prominent under warmer developmental conditions, whereas HN showed broader changes in stress- and metabolism-associated pathways under high temperature. Overall, this study provides a comparative transcriptomic analysis of two field-collected S. exigua populations under different developmental temperature regimes and identifies RNA-seq-based molecular response patterns associated with population-specific thermal response profiles.

Animals

Analysis of differentially expressed genes in schizophrenia based on bioinformatics and corresponding mRNA expression levels.

OBJECTIVE: This study aimed to use bioinformatics analysis to identify differentially expressed genes (DEGs) involved in the pathogenesis of schizophrenia and validate their mRNA expression levels through real-time quantitative PCR (qPCR). MATERIAL/METHODS: Datasets from the publicly available Gene Expression Omnibus (GEO) database were analyzed using R software to identify DEGs. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, were conducted. A protein-protein interaction (PPI) network was constructed using Cytoscape software to identify key genes with notable expression changes. The expression levels of these key genes were subsequently validated in schizophrenia patients using qPCR to assess potential susceptibility genes. RESULTS: In total, 813 DEGs were identified, with six key genes highlighted through GO analysis and PPI network screening. Among these, HDAC1, UBA52, and FYN demonstrated statistically significant differences in mRNA expression between schizophrenia patients and healthy controls (P&#xa0;<&#xa0;0.05). CONCLUSIONS: This study identified several DEGs potentially linked to the pathogenesis of schizophrenia, suggesting that HDAC1, UBA52, and FYN could serve as candidate susceptibility genes and diagnostic biomarkers. These findings provide new insights and directions for future schizophrenia research.

Humans

DiaReport: reproducible workflow for differential expression analysis and interactive reporting in DIA-based proteomics.

MOTIVATION: Data-independent acquisition (DIA) has become the preferred data acquisition method for mass spectrometry-based proteomics, yet, reproducible workflows for differential expression (DE) analysis and results reporting remain limited. We present DiaReport, an R package that performs precursor- and protein-level DE analysis from DIA-NN output using MSqRob and QFeatures, while generating high-quality, interactive HTML reports through Quarto. DiaReport integrates precursor data, filtering of missing values, normalization, protein summarization and statistical modeling within a single function, supporting both simple pairwise as well as complex experimental designs. The package provides structured outputs and configuration files to ensure computational reproducibility across different studies. To accommodate diverse research needs, DiaReport includes multiple reporting templates tailored to different proteomic applications. Applying DiaReport to an extracellular vesicle (EV) proteomics dataset demonstrates its ability to efficiently analyze DIA data and provide rapid insights into sample quality and protein level differences. AVAILABILITY: DiaReport is an open-source R package available at https://github.com/Gevaert-Lab/diareport (DOI: 10.5281/zenodo.20120604). The package is platform-independent and distributed under the MIT license. Reports are generated using Quarto and require only standard R dependencies. Detailed documentation, installation guides and usage vignettes are provided within the repository. The interactive HTML reports discussed in this study, including the UPS2 benchmark and EV case study, are archived on Zenodo (10.5281/zenodo.20122506 and 10.5281/zenodo.20123378).

Proteomics

Transcriptome analysis reveals that PRV XJ delgE/gI/TK protects against intestinal damage in nose-dropping-infected mice by regulating ECM-ITGA/ITGB-P-FAK.

Pseudorabies virus (PRV) is an ideal model for mechanistic investigations into &#x3b1;-herpesvirus. The neurotropism and latent infection of PRV have been extensively studied. Apart from neurological symptoms, diarrhea caused by PRV infection is also an essential cause of mortality in newborn and weaned piglets. However, little research has been done on PRV invasion of the gut. To fill this gap, a nasal drip PRV-infection mouse model was developed, consisting of three groups: the challenged group (Group A), the immunization-challenged group (Group B), and a mock group (Group C). The results showed that immunization with PRV XJ delgE/gI/TK successfully prevented intestinal damage caused by PRV drop-nose infection. Subsequently, intestines were collected for transcriptional analysis. Differentially expressed genes analysis revealed that PRV XJ delgE/gI/TK was effective in reducing the organismal intestinal transcriptional activity caused by PRV. The Group A vs Group C and Group A vs Group B had similar Kyoto Encyclopedia of Genes and Genomes (KEGG)-enriched signaling pathways and the differentially expressed genes were primarily enriched in pathways, such as cell adhesion molecules, focal adhesion kinase, and actin cytoskeleton regulation. Notably, transcriptome analysis indicated that genes associated with the focal adhesion kinase (FAK) signaling pathway (ECM-ITGA/ITGB-p-FAK) were significantly more highly expressed in Group A than in Group B and Group C. The results of quantitative real-time PCR (RT-qPCR) and western blotting were consistent with KEGG analysis. Therefore, we hypothesized that PRV promotes self-infection through activation of the ECM-ITGA/ITGB-p-FAK signaling pathway and that PRV XJ delgE/gI/TK immunization could attenuate the intestinal damage caused by PRV by inhibiting the activation of this pathway.IMPORTANCEPseudorabies virus (PRV) poses a significant threat to the swine industry and public health due to its ability to infect multiple species, including humans, leading to substantial economic losses and potential health risks. This study addresses a critical gap in understanding the impact of PRV infection on the gut, which has been less explored compared to its neurological effects. By developing a drip-nose PRV-infection mouse model, the research indicated that PRV might promote self-infection through activation of the ECM-ITGA/ITGB-p-FAK signaling pathway, and PRV XJ delgE/gI/TK immunization effectively prevents intestinal damage by significantly reducing the expression of genes in the ECM-ITGA/ITGB-p-FAK signaling pathway. The research has important implications for the swine industry and public health by contributing to the development of better vaccines and treatments, ultimately helping to control PRV and prevent its cross-species transmission.

Animals

Atlas-level single-cell integration and clustering-free differential expression analysis with GEDI 2.0.

MOTIVATION: GEDI is a generative framework for multi-sample, multi-condition single-cell analysis that performs batch correction, latent representation learning, and clustering-free differential expression within a unified model. However, the original implementation suffered from prohibitive memory use and runtime, preventing its application to modern atlas-scale datasets. RESULTS: We present GEDI 2.0, a complete high-performance reimplementation featuring a standalone C++ computational core with pre-allocated workspaces, strict sparse-matrix preservation, optimized BLAS routines, and multi-threaded block-coordinate descent. Across extensive benchmarks spanning up to 500 000 cells and 10 000 features, GEDI 2.0 achieves 40%-63.6% mean reduction in peak memory, 2.98&#xd7; mean single-threaded speedups, and up to 11.5&#xd7; acceleration with parallel execution, while maintaining full numerical equivalence to the original method. These improvements enable GEDI 2.0 to analyze million-cell datasets, a scale not achievable with the legacy implementation. GEDI 2.0 provides R and Python interfaces and seamless interoperability with common single-cell workflows. AVAILABILITY AND IMPLEMENTATION: Source code, documentation, reproducible codebase, and tutorials are available at https://github.com/csglab/gedi2.

Single-Cell Analysis

damidBind: an R/bioconductor package for differential DamID analysis and data exploration.

SUMMARY: DamID, and its cell-type specific adaptations, including Targeted DamID (TaDa) and Chromatin Accessibility TaDa (CATaDa), are now widely-adopted as techniques for the genome-wide profiling of DNA binding proteins. Despite this popularity, no dedicated software solution exists for identifying differentially bound or accessible loci, or differentially transcribed genes, between cell types using DamID. The R/Bioconductor package damidBind provides these functions, allowing an end-user to move from processed binding profiles to identifying differentially-bound loci in a reproducible, statistically appropriate and straightforward workflow. AVAILABILITY AND IMPLEMENTATION: damidBind is an open-source R/Bioconductor package and freely available from Bioconductor at https://bioconductor.org/packages/damidBind/, and from GitHub at https://github.com/marshall-lab/damidBind. It is released under the GPLv3 licence.

Software

Comprehensive Analysis of Differentially Expressed Genes and Immune Infiltration in Burn Injury: Key Biomarkers and Pathways.

BACKGROUND: Burn injuries trigger complex immune responses and gene expression changes, impacting wound healing and systemic inflammation. Understanding these changes is crucial for identifying biomarkers and therapeutic targets. METHODS: We analyzed two gene expression omnibus datasets (wound tissue [GSE8056] and blood [GSE37069]) to identify differentially expressed genes (DEGs) in burn injury samples versus controls. Immune cell proportions were assessed using CIBERSORT. Functional enrichment analyses (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes) and protein-protein interaction networks were constructed to identify key genes and pathways. RESULTS: We identified 1170 upregulated and 1227 downregulated DEGs. Gene Ontology analysis revealed enrichment in neutrophil activation, inflammatory response, and extracellular matrix organization. Kyoto Encyclopedia of Genes and Genomes analysis highlighted cytokine-cytokine receptor interaction, TNF, and IL-17 signaling pathways. Immune infiltration analysis showed significant changes in neutrophils, macrophages (M1/M2), and T-cell subsets. Protein-protein interaction network analysis identified five hub genes: JUN, STAT1, Bcl2, MMP9, and TLR2. CONCLUSIONS: This study provides a comprehensive bioinformatic analysis of gene expression and immune responses in burn injuries. The identified DEGs, hub genes, and pathways offer insights into the immune response mechanisms and suggest potential targets for diagnostic and therapeutic interventions in burn injury management.

Burns

Non-parametric differential methylation analysis characterizes histotype-specific promoter regions in epithelial ovarian cancer.

Epithelial ovarian cancer (EOC) is a heterogenous disease with frequent late-stage diagnosis and high mortality rates, for which no reliable screening tests exist. In recent years, epigenetic biomarkers in the form of DNA methylation in CpG-rich regions have gained increased attention in the scientific community due to their robust nature and accessibility, allowing for diagnosis without the need for invasive surgery. In this study, we investigated the aberrant methylation of promoter regions in early stage EOC through non-parametric methods, with the purpose of characterizing candidate epigenetic biomarkers. The approach was used on a cohort of early stage EOC samples, and results were compared to existing programs for differential methylation. Significant regions were then used to construct a CpG panel for stratifying EOC histotypes through predictive classification in external data. Identified promoter regions were highly reproducible across cohorts, and the constructed CpG model stratified histotypes in external cohorts through predictive classification. Comparisons against other DMP and DMR callers showed a degree of homogeneity between results but also revealed promoter regions that were overlooked despite clear signs of aberrant methylation. Finally, EOC histotypes were found to differ in their methylation distribution types, and results indicate that methods sensitive to non-normally distributed data may be poorly suited to compare groups with different distribution types. The non-parametric approach identified aberrantly methylated promoter regions that were highly reproducible across cohorts. Results from predictive classification indicate that these regions may be useful for the purpose of EOC histotype stratification.

Humans

Comprehensive analysis of differentially expressed mRNAs, lncRNAs, and miRNAs involved in ovarian differentiation and development in Qihe gibel carp (Carassius gibelio var. Qihe).

Qihe gibel carp (Carassius gibelio var. Qihe) exhibits diverse reproductive modes including gynogenesis and sexual reproduction, yet the molecular mechanisms of ovarian differentiation remain poorly understood. Ovarian tissues at 20, 30, and 60&#xa0;days after hatching (dah), representing key stages covering early ovarian differentiation and primary oocyte growth, were subjected to whole-transcriptome sequencing. A total of 27,259 mRNAs, 2622 lncRNAs, and 2467 miRNAs were differentially expressed. Cell cycle, transcription, translation, and DNA replication pathways were significantly upregulated from 20 to 60 dah. Oocyte meiosis was enriched from 20 and 30 dah, whereas metabolic pathways (lipid, carbohydrate, and nucleotide metabolism) were enriched from 30 to 60 dah, indicating sequential progression from meiosis initiation to primary oocyte growth with nutrient synthesis. Hub lncRNAs and key ceRNA networks (e.g., MSTRG.28669.5-miR-221-ccnb2) were identified. This study provides the first comprehensive characterization of ncRNA-mediated regulation and ceRNA networks during ovarian development in Qihe gibel carp, establishing a foundation for understanding ovarian differentiation in this species.

Animals

Genomic analysis of differentiation and demography of the formerly conspecific agile (Dipodomys agilis) and Dulzura (D. simulans) kangaroo rats.

Karyotype variation within Pacific kangaroo rat Dipodomys agilis motivated its division in 1997 into the agile kangaroo rat (AKR, D. agilis, 2N&#x2009;=&#x2009;62) in the north of its range in California, and Dulzura kangaroo rat (DKR, D. simulans, 2N&#x2009;=&#x2009;60) to the south, with a suspected sympatric zone south of the San Gabriel and San Bernardino Mountains. This division was supported by our whole genome sequencing that sampled a ~120&#x2009;km transect from north of the mountains to SW Riverside County. The taxa showed marked genetic differentiation, with no evidence of hybridization or sympatry. AKR was found at the southern edge of the mountains, precluding the mountain barrier driving isolation, suggesting ecological separation linked to habitat differences between the mountains and the arid area to the south. Adding four additional Dipodomys species, we estimated genetic divergence times in the genus back to &#x223c;3.5&#x2009;mya. AKR and DKR diverged from D. stephensi &#x223c;1.7&#x2009;mya, and from each other &#x223c;0.5&#x2009;mya, when their joint effective population size (Ne) was ~100,000. After separation, DKR's Ne declined to ~20,000, while AKR's was little changed. More recently their Ne converged at ~50,000. Runs of homozygosity were longer in AKR, indicating a smaller neighborhood size, which may have promoted the karyotype change; however, nucleotide diversity was higher in AKR, but both had levels typical for rodents, indicating neither experienced recent bottlenecks. These patterns provide a baseline for any future conservation efforts. More generally, this study shows how a detailed genomic study can resolve taxonomic and demographic questions among morphologically indistinguishable taxa.

Animals

Anoikis classification of lung squamous cell carcinoma reveals correlation with clinical prognosis and immune characteristics.

BACKGROUND: Anoikis is a new mode of cell death that has been shown to correlate significantly with tumors. However, the clinical prognostic significance of anoikis in lung squamous cell carcinoma (LUSC) remains poorly studied. METHODS: The differentially expressed ARGs and candidate genes were selected by the differential analysis to construct a predictive model. Independent prognostic gene was determined by Cox and LASSO analysis and we used the HCC95 and NCI H520 cell line to verify the gene function. We used the data from TCGA, GEO, GeneCards, and Harmonizome databases to analyze the immune microenvironment, functional enrichment, and drug sensitivity analysis. RESULTS: We identified 717 differentially expressed and selected 3 ARGs (FADD, SNAI1, and BAG4) to construct a predictive model. We found that SNAI1 is an independent prognostic gene and confirmed that knocking out the SNAI1 inhibited the HCC95/NCI H520 cell proliferation. We used single-sample gene-set enrichment analysis (ssGSEA) to evaluate the immune infiltration based on the 3 ARG expression levels. We constructed a risk score and provided a visual representation of the prophetic implications of the ARGs-based signature through a nomogram. We found 15 susceptible drugs in the high-risk group and 15 sensitive drugs in the low-risk group by the drug sensitivity analysis. CONCLUSION: We used ARGs to construct a prognosis model for LUSC that can accurately predict the prognosis of LUSC patients. ARGs, especially SNAI1, play an essential role in developing LUSC. These findings could provide individualized treatment plans and new research ideas for LUSC patients.

Humans

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

Identification of potential biomarkers and mechanisms for keloid disorder based on comprehensive bioinformatics analysis and machine learning algorithms.

BACKGROUND: Keloid disorder (KD) encompasses a spectrum of fibroproliferative dermal conditions, the pathogenesis remains complex and incompletely understood. This study sought to identify biomarkers and potential therapeutic targets for KD through an integrative bioinformatics approach and machine learning analysis of RNA sequencing data. METHODS: RNA sequencing was performed on skin tissue samples from 13 patients with KD and 14 healthy controls. Using weighted gene co-expression network analysis and differential expression analysis revealed differentially expressed key module genes, and the CytoHubba plugin identified candidate genes. Subsequently analyzed using least absolute shrinkage and selection operator (LASSO) and support vector machine recursive feature elimination (SVM-RFE) methods to pinpoint feature genes associated with KD. Following this, biomarkers were determined through expression level validation, enrichment analysis, and immune infiltration analysis. RESULTS: A total of 420 differentially expressed key module genes were identified, and the top 10 genes with DMNC values were selected as candidate genes. Five feature genes were selected through LASSO and SVM-RFE, with NID2, MFAP2, COL8A1, and P4HA3 showing significant expression differences between KD and control samples, along with consistent expression patterns across datasets, identified as potential biomarkers. These four biomarkers were proved to possess high diagnostic potential, and they were found to exhibit significant positive correlations with one another. Functional enrichment analysis indicated that the primary KEGG pathways associated with these biomarkers included "steroid hormone biosynthesis" and "cytokine-cytokine receptor interaction." Moreover, immune infiltration analysis revealed that the four biomarkers were negatively correlated with type 17 T helper cells and positively correlated with 15 immune cell types, including activated B cells and central memory CD4 T cells. CONCLUSION: In conclusion, NID2, MFAP2, COL8A1, and P4HA3 were identified as key biomarkers for KD, offering new avenues for more targeted and effective diagnostic and therapeutic strategies for managing this condition.

Humans