Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan–Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans

On the use of machine learning to identify topological rules in the packing of beta-strands.

The machine learning program GOLEM was applied to discover topological rules in the packing of beta-sheets in alpha/beta-domain proteins. Rules (constraints) were determined for four features of beta-sheet packing: (i) whether a beta-strand is at an edge; (ii) whether two consecutive beta-strands pack parallel or anti-parallel; (iii) whether two beta-strands pack adjacently; and (iv) the winding direction of two consecutive beta-strands. Rules were found with high predictive accuracy and coverage. The errors were generally associated with complications in domain folds, especially in one doubly would domains. Investigation of the rules revealed interesting patterns, some of which were known previously, others that are novel. Novel features include (i) the relationship between pairs of sequential strands is in general one of decreasing size; (ii) more sequential pairs of strands wind in the direction out than in; and (iii) it takes a larger alteration in hydrophobicity to change a strand from winding in the direction out than in. These patterns in the data may be the result of folding pathways in the domains. The rules found are of predictive value and could be used in the combinatorial prediction of protein structure, or as a general test of model structures, e.g. those produced by threading. We conclude that machine learning has a useful role in the analysis of protein structures.

Amino Acid Sequence

Reconstructing muscle activation during normal walking: a comparison of symbolic and connectionist machine learning techniques.

One symbolic (rule-based inductive learning) and one connectionist (neural network) machine learning technique were used to reconstruct muscle activation patterns from kinematic data measured during normal human walking at several speeds. The activation patterns (or desired outputs) consisted of surface electromyographic (EMG) signals from the semitendinosus and vastus medialis muscles. The inputs consisted of flexion and extension angles measured at the hip and knee of the ipsilateral leg, their first and second derivatives, and bilateral foot contact information. The training set consisted of data from six trials, at two different speeds. The testing set consisted of data from two additional trials (one at each speed), which were not in the training set. It was possible to reconstruct the muscular activation at both speeds using both techniques. Timing of the reconstructed signals was accurate. The integrated value of the activation bursts was less accurate. The neural network gave a continuous output, whereas the rule-based inductive learning rule tree gave a quantised activation level. The advantage of rule-based inductive learning was that the rules used were both explicit and comprehensible, whilst the rules used by the neural network were implicit within its structure and not easily comprehended. The neural network was able to reconstruct the activation patterns of both muscles from one network, whereas two separate rule sets were needed for the rule-based technique. It is concluded that machine learning techniques, in comparison to explicit inverse muscular skeletal models, show good promise in modelling nearly cyclic movements such as locomotion at varying walking speeds.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape

MULTIPREVENT: Integrated screening for smoking-related multimorbidity using low-dose chest computed tomography.

OBJECTIVES: Tobacco consumption, combined with individual genetic predispositions, contributes to an age-dependent risk not only for lung cancer but also for other non-communicable diseases (NCDs) such as cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), osteoporosis, and diabetes. The MULTIPREVENT project aims to validate whether low-dose computed tomography (LDCT) of the chest, combined with simple biomarkers, functional tests, and genomic profiling, can serve as an effective tool for comprehensive health assessment and risk prediction of multimorbidity in adults. STUDY DESIGN: The study is based on a prospective epidemiological design involving 3000 participants from the MOLTEST-BIS lung cancer screening cohort (2016-2018). These participants, aged 50-79 years (during MOLTEST-BIS) and with a smoking history of at least 30 pack-years, will undergo two follow-up assessments in 2025-2027 and 2030-2032. METHODS: Each follow-up includes LDCT, spirometry, standardized blood pressure measurement, anthropometric evaluation, biomarker assessment (lipid profile, lipoprotein(a), glycated haemoglobin), and health-related questionnaires. Genetic profiling will be performed using the Illumina Infinium Global Screening Arrays approach to identify inherited predispositions to major NCDs. All data, clinical, imaging (including radiomics), molecular, and genetic, will be integrated through machine learning algorithms to develop AI-based risk prediction models. RESULTS: The MULTIPREVENT study is expected to generate a wide range of scientific, clinical, and infrastructural results that will serve as a foundation for future public health initiatives in integrated prevention. CONCLUSIONS: By linking imaging and biochemical markers, genetic susceptibility, and clinical parameters within a longitudinal design, MULTIPREVENT will establish data-driven, AI-supported prevention strategies aimed at reducing morbidity and mortality among adults exposed to tobacco. The project will also serve as a model for population-based multimorbidity prevention programs.

Humans

Machine learning approaches for the prediction of signal peptides and other protein sorting signals.

Prediction of protein sorting signals from the sequence of amino acids has great importance in the field of proteomics today. Recently, the growth of protein databases, combined with machine learning approaches, such as neural networks and hidden Markov models, have made it possible to achieve a level of reliability where practical use in, for example automatic database annotation is feasible. In this review, we concentrate on the present status and future perspectives of SignalP, our neural network-based method for prediction of the most well-known sorting signal: the secretory signal peptide. We discuss the problems associated with the use of SignalP on genomic sequences, showing that signal peptide prediction will improve further if integrated with predictions of start codons and transmembrane helices. As a step towards this goal, a hidden Markov model version of SignalP has been developed, making it possible to discriminate between cleaved signal peptides and uncleaved signal anchors. Furthermore, we show how SignalP can be used to characterize putative signal peptides from an archaeon, Methanococcus jannaschii. Finally, we briefly review a few methods for predicting other protein sorting signals and discuss the future of protein sorting prediction in general.

Algorithms

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article

How to interpret an anonymous bacterial genome: machine learning approach to gene identification.

In this report we address the problem of accurate statistical modeling of DNA sequences, either coding or noncoding, for a bacterial species whose genome (or a large portion) was sequenced but not yet characterized experimentally. Availability of these models is critical for successful solution of the genome annotation task by statistical methods of gene finding. We present the method, GeneMark-Genesis, which learns the parameters of Markov models of protein-coding and noncoding regions from anonymous bacterial genomic sequence. These models are subsequently used in the GeneMark and GeneMark.hmm gene-finding programs. Although there is basically one model of a noncoding region for a given genome, several models of protein-coding region are automatically obtained by GeneMark-Genesis. The diversity of protein-coding models reflects the diversity of oligonucleotide compositions, particularly the diversity of codon usage strategies observed in genes from one and the same genome. In the simplest and the most important case, there are just two gene models-typical and atypical ones. We show that the atypical model allows one to predict genes that escape identification by the typical model. Many genes predicted by the atypical model appear to be horizontally transferred genes. The early versions of GeneMark-Genesis were used for annotating the genomes of Methanoccocus jannaschii and Helicobacter pylori. We report the results of accuracy testing of the full-scale version of GeneMark-Genesis on 10 completely sequenced bacterial genomes. Interestingly, the GeneMark.hmm program that employed the typical and atypical models defined by GeneMark-Genesis was able to predict 683 new atypical genes with 176 of them confirmed by similarity search.

Algorithms

Sparse deconvolution of cell type medleys in spatial transcriptomics.

Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling However, current deconvolution models aim to estimate the proportions of distinct cell types in each spatial transcriptomics spot by integrating reference single-cell data. These models often assume strong overlap between the reference and spatial datasets, neglecting biology-grounded constraints such as sparsity and cell-type variations, as well as technical sparsity. As a result, these methods rely on over-permissive algorithms that ignore given constraints leading to inaccurate predictions, particularly in heterogeneous or unmatched datasets. We introduce Weight-Induced Sparse Regression (WISpR), a machine learning algorithm that integrates spot-specific hyperparameters and sparsity-driven modeling. Unlike conventional approaches that neglect biology-grounded constraints, WISpR accurately predicts cell-type distributions while preserving biological coherence, i.e., spatially and functionally consistent cell-type localization, even in unmatched datasets. Benchmarking against five alternative methods across ten datasets, WISpR consistently outperformed competitors and predicted cellular landscapes in both normal and cancerous tissues. By leveraging sparse cell-type arrangements, WISpR provides biologically informed, high-resolution cellular maps. Its ability to decode tissue organization in both healthy and diseased states highlights WISpR's practical utility for spatial transcriptomics, particularly in challenging settings involving noise, sparsity, or reference mismatches.

Humans

Prediction of primate splice junction gene sequences with a cooperative knowledge acquisition system.

We propose a cooperative conceptual modelling environment in which two agents interact: the machine and the human expert. The former is able to extract knowledge from data using a symbolic-numeric machine learning system, and the latter is able to control the learning process by accepting and validating the machine results, or by criticizing those results or the explanation that the system produces on them. The improvement of the conceptual modelling relies on the cooperation between the two agents. Results obtained with our method on prediction of primate splice junctions sites in genetic sequences are far better than those reported in the literature with other symbolic machine learning systems, and are as better as those obtained with some artificial neural networks methods reported at present. But in opposite to neural networks which lack of argumentation, our system provides the user a plausible explanation of its prediction.

Algorithms

Integrative multi-omics and machine learning identify the SPI1-METTL16-PLIN4 axis as a candidate driver of steatosis in HepG2 cells.

BACKGROUND: Non-alcoholic fatty liver disease (NAFLD) is a prevalent metabolic disorder with limited therapeutic options. This study aimed to identify potential regulators and explore their functional roles in a cellular model of NAFLD. METHODS: WGCNA was performed on the hepatic transcriptomic dataset GSE126848 (31 NAFLD vs. 26 controls), followed by integration with serum proteomic data from 12 NAFLD patients and 12 healthy controls. Hub genes were prioritized using three machine learning algorithms. Functional validation was conducted in a HepG2 cellular steatosis model induced by high fructose (3.2&#x202f;g/L) and oleic acid (400&#x202f;&#x3bc;M) for 48&#x202f;h. Lipid accumulation was assessed by Oil Red O staining and triglyceride/total cholesterol measurement. Inflammation was evaluated by TNF-&#x3b1; and IL-6 secretion (ELISA), and oxidative stress by ROS levels (flow cytometry). The binding interaction between METTL16 and PLIN4 mRNA was validated by RNA immunoprecipitation (RIP)-quantitative PCR. METTL16-mediated m6A modification of PLIN4 was assessed by Methylated RIP (MeRIP)-quantitative PCR. Transcriptional regulation of METTL16 by SPI1 was examined by chromatin immunoprecipitation (ChIP) and dual-luciferase reporter assays. RESULTS: Integrative analysis identified PLIN4 as a core hub gene. PLIN4 was upregulated in the HepG2 steatosis model (P&#x202f;<&#x202f;0.001). PLIN4 knockdown alleviated lipid droplet accumulation (P&#x202f;<&#x202f;0.001), reduced TNF-&#x3b1; and IL-6 secretion (P&#x202f;<&#x202f;0.01), and decreased ROS levels (P&#x202f;<&#x202f;0.001) in fructose/oleic acid-treated HepG2 cells. Mechanistically, METTL16 mediated its m6A modification to enhance PLIN4 mRNA stability. Furthermore, SPI1 was found to transcriptionally activate METTL16 by binding to its promoter (P&#x202f;<&#x202f;0.001). PLIN4 re-expression partially reversed the protective effects of SPI1 knockdown on lipid accumulation (P&#x202f;=&#x202f;0.01), inflammation (P&#x202f;<&#x202f;0.05), and oxidative stress (P&#x202f;<&#x202f;0.001). CONCLUSION: This study identifies the SPI1/METTL16/PLIN4 axis as a potential regulatory mechanism contributing to in vitro steatosis, inflammation, and oxidative stress in steatotic HepG2 cells.

Humans

Evaluation of automatic knowledge acquisition techniques in the diagnosis of acute abdominal pain. Acute Abdominal Pain Study Group.

Clinical diagnosis in acute abdominal pain is still a major problem. Computer-aided diagnosis offers some help; however, existing systems still produce high error rates. We therefore tested machine learning techniques in order to improve standard statistical systems. The investigation was based on a prospective clinical database with 1254 cases, 46 diagnostic parameters and 15 diagnoses. Independence Bayes and the automatic rule induction techniques ID3, NewId, PRISM, CN2, C4.5 and ITRULE were trained with 839 cases and separately tested on 415 cases. No major differences in overall accuracy were observed (43-48%), except for NewId, which was below the average. Between the different techniques some similarities were found, but also considerable differences with respect to specific diagnoses. Machine learning techniques did not improve the results of the standard model Independence Bayes. Problem dimensionality, sample size and model complexity are major factors influencing diagnostic accuracy in computer-aided diagnosis of acute abdominal pain.

Abdominal Pain

Major depletion of insulin sensitivity-associated taxa in the gut microbiome of persons living with HIV controlled by antiretroviral drugs.

BACKGROUND: Persons living with HIV (PWH) harbor an altered gut microbiome (higher abundance of Prevotella and lower abundance of Bacillota and Ruminococcus lineages) compared to non-infected individuals. Some of these alterations are linked to sexual preference and others to the HIV infection. The relationship between these lineages and metabolic alterations, often present in aging PWH, has been poorly investigated. METHODS: In this study, we compared fecal metagenomes of 25 antiretroviral-treatment (ART)-controlled PWH to three independent control groups of 25 non-infected matched individuals by means of univariate analyses and machine learning methods. Moreover, we used two external datasets to validate predictive models of PWH classification. Next, we searched for associations between clinical and biological metabolic parameters with taxonomic and functional microbiome profiles. Finally, we compare the gut microbiome in 7 PWH after a 17-week ART switch to raltegravir/maraviroc. RESULTS: Three major enterotypes (Prevotella, Bacteroides and Ruminococcaceae) were present in all groups. The first Prevotella enterotype was enriched in PWH, with several of characteristic lineages associated with poor metabolic profiles (low HDL and adiponectin, high insulin resistance (HOMA-IR)). Conversely butyrate-producing lineages were markedly depleted in PWH independently of sexual preference and were associated with a better metabolic profile (higher HDL and adiponectin and lower HOMA-IR). Accordingly with the worst metabolic status of PWH, butyrate production and amino-acid degradation modules were associated with high HDL and adiponectin and low HOMA-IR. Random Forest models trained to classify PWH vs. control on taxonomic abundances displayed high generalization performance on two external holdout datasets (ROC AUC of 80-82%). Finally, no significant alterations in microbiome composition were observed after switching to raltegravir/maraviroc. CONCLUSION: High resolution metagenomic analyses revealed major differences in the gut microbiome of ART-controlled PWH when compared with three independent matched cohorts of controls. The observed marked insulin resistance could result both from enrichment in Prevotella lineages, and from the depletion in species producing butyrate and involved into amino-acid degradation, which depletion is linked with the HIV infection.

Humans

Stratifying lung adenocarcinoma: a novel prognostic model based on mitochondrial outer membrane permeabilization activity.

UNLABELLED: Mitochondrial outer membrane permeabilization (MOMP) is a core apoptotic regulatory event that dictates mitochondrial integrity, where full activation drives cell death and sublethal dysregulation contributes to tumor genomic instability. We used the Cancer Genome Atlas lung adenocarcinoma cohort (TCGA-LUAD) as the training cohort and the Gene Expression Omnibus dataset GSE42127 as the validation cohort to identify prognostic genes related to MOMP activity in lung adenocarcinoma (LUAD) and to evaluate their potential biological significance. By intersecting MOMP-related genes with differentially expressed genes, combined with survival analysis, Mendelian randomization analysis, and 101 machine-learning algorithm combinations, seven prognostic genes, namely BIRC5, PSMD11, TNFRSF13C, YWHAZ, YWHAG, CYCS, and LTB, were identified. Next, an optimal prognostic model was constructed based on the gradient boosting machine (GBM) algorithm. Based on the risk score, LUAD patients were stratified into high- and low-risk groups, and patients in the high-risk group exhibited poorer overall survival in both the training and validation cohorts. Furthermore, a nomogram integrating the risk score and clinicopathological factors was developed and showed favorable predictive performance for 1-, 3-, and 5-year survival. Meanwhile, functional and immune analyses revealed that the high-risk group was enriched in DNA replication-related pathways and demonstrated a higher tumor mutation burden (TMB). Correlation analysis indicated that TNFRSF13C was positively correlated with activated B cells, whereas BIRC5 was negatively correlated with eosinophils, suggesting that MOMP-related genes might be involved in remodeling the immune microenvironment of LUAD. Drug sensitivity analysis showed differences in predicted half-maximal inhibitory concentration (IC50) values between the risk groups, suggesting the potential value of this model in assisting therapeutic stratification. Single-cell RNA sequencing (scRNA-seq) further identified T lymphocytes as a key cell type, with numerous prognostic genes exhibiting differential expression in T cells or dynamic changes during differentiation. We suggest that the MOMP-related signature established in this study may provide a reference for prognostic stratification in LUAD and offers candidate prognostic genes for subsequent experimental and clinical validation. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05058-6.

Lung adenocarcinoma