Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

A Risk Score for Polycystic Ovary Syndrome Based on Meta-Analysis and Machine Learning of Gut Microbiota Signatures.

Polycystic Ovary Syndrome (PCOS) is a prevalent endocrine and metabolic disorder among reproductive-age women, in which emerging evidence suggests a substantial role played by the gut microbiota. To comprehensively evaluate gut microbiota alterations in PCOS and identify microbial biomarkers through integrated analysis, a systematic search of PubMed, Web of Science, and Embase was conducted for studies employing 16S rRNA gene sequencing of fecal samples from PCOS cohorts. Ten eligible PCOS cohorts, comprising 858 individuals, were included in the study, from which a risk score was derived using a 20-gene gut microbial signature associated with PCOS. Meta-analysis at the genus level identified that Subdoligranulum, NK4A214_group, and Collinsella significantly decreased, and Bacteroides increased in PCOS across multiple cohorts. Machine learning analysis identified a 20-genus microbial signature using the least absolute shrinkage and selection operator (LASSO) method, which was used to construct a risk score with an AUC of 0.835 in diagnosis prediction. Network analysis further identified Negativibacillus and Lachnospiraceae_UCG_010 as potential driver microbes in PCOS. The analysis in this study highlights key alterations in the gut microbiota across PCOS cohorts. The identified gut microbial signature and derived LASSO-based risk model offer novel insights and a potential tool for PCOS diagnosis.

Polycystic Ovary Syndrome↗

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans↗

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article↗

Acceptance of rules generated by machine learning among medical experts.

OBJECTIVES: The aim was to evaluate the potential for monotonicity constraints to bias machine learning systems to learn rules that were both accurate and meaningful. METHODS: Two data sets, taken from problems as diverse as screening for dementia and assessing the risk of mental retardation, were collected and a rule learning system, with and without monotonicity constraints, was run on each. The rules were shown to experts, who were asked how willing they would be to use such rules in practice. The accuracy of the rules was also evaluated. RESULTS: Rules learned with monotonicity constraints were at least as accurate as rules learned without such constraints. Experts were, on average, more willing to use the rules learned with the monotonicity constraints. CONCLUSIONS: The analysis of medical databases has the potential of improving patient outcomes and/or lowering the cost of health care delivery. Various techniques, from statistics, pattern recognition, machine learning, and neural networks, have been proposed to "mine" this data by uncovering patterns that may be used to guide decision making. This study suggests cognitive factors make learned models coherent and, therefore, credible to experts. One factor that influences the acceptance of learned models is consistency with existing medical knowledge.

Alzheimer Disease↗

Machine learning of motor vehicle accident categories from narrative data.

Bayesian inferencing as a machine learning technique was evaluated for identifying pre-crash activity and crash type from accident narratives describing 3,686 motor vehicle crashes. It was hypothesized that a Bayesian model could learn from a computer search for 63 keywords related to accident categories. Learning was described in terms of the ability to accurately classify previously unclassifiable narratives not containing the original keywords. When narratives contained keywords, the results obtained using both the Bayesian model and keyword search corresponded closely to expert ratings (P(detection) > or = 0.9, and P (false positive) < or = 0.05). For narratives not containing keywords, when the threshold used by the Bayesian model was varied between p > 0.5 and p > 0.9, the overall probability of detecting a category assigned by the expert varied between 67% and 12%. False positives correspondingly varied between 32% and 3%. These latter results demonstrated that the Bayesian system learned from the results of the keyword searches.

Accidents, Traffic↗

Phylogenetic Methods Meet Deep Learning.

Deep learning (DL) has been widely used in various scientific fields, but its integration into phylogenetics has been slower, primarily due to the complex nature of phylogenetic data. The studies that apply DL to sequencing data often limit analyses to four-taxon trees. Many of these studies serve as "proof of principle" and perform similarly to traditional phylogeny reconstruction methods. New ways of using training data, such as encoding with compact bijective ladderized vectors or transformers, enable the handling of much larger trees and genomic data sets. This short perspective focuses on the application of DL in phylogenetics, introducing prevalent DL architectures. We highlight potential problems in the field by discussing the risks of using simulation-based training data and emphasize the importance of reproducibility and robustness in computational estimates. Finally, we explore promising research areas, including the combination of phylogenetics and population genetics in DL, the analysis of neighbor dependencies, and the potential to significantly reduce computational cost compared to traditional methods. This perspective illustrates the potential of DL in complementing traditional phylogeny reconstruction methods and aiding the advancement of phylogenetic analysis, especially in performing computationally demanding tasks such as model selection or estimating branch support values.

Humans↗

Proteomic signatures and predictive modeling of cadmium-associated anxiety in middle-aged and elderly populations: an environmental exposure association study.

BACKGROUND: Emerging evidence implicates environmental contaminants such as cadmium (Cd) as modifiable risk factors for anxiety. Despite growing recognition of heavy metal toxicity in neuropsychiatric disorders, the molecular mechanisms linking environmental exposure to anxiety pathogenesis remain poorly understood. METHODS: Based on the established cohort of individuals with cognitive impairment in cadmium-contaminated areas, this cross-sectional association study enrolled 50 middle-aged and elderly hospitalized patients from these regions, adhering to the STROBE guidelines. Blood concentrations of cadmium (Cd), lead (Pb), and mercury (Hg) were analyzed in relation to anxiety severity assessed via the Hamilton Anxiety Rating Scale (HAMA). Plasma proteomic profiling was performed using data-independent acquisition (DIA) quantitative technology with an LC-MS/MS platform (timsTOF Pro, Bruker Daltonics), systematically characterizing 2,531 proteins across all samples. Machine learning techniques, specifically XGBoost and LASSO, were employed to identify biomarkers that were subsequently validated through mediation analysis and animal experiments, allowing for the screening of key protein signatures. Finally, clinical variables were integrated to construct a comprehensive model, which was then thoroughly evaluated. RESULTS: Anxious individuals exhibited significantly higher blood Cd levels than controls (&#x3b2;&#x2009;=&#x2009;0.50, 95% CI: 0.07-0.93, p&#x2009;<&#x2009;0.01), with anxiety positively correlating with depression (r&#x2009;=&#x2009;0.62, p&#x2009;=&#x2009;0.003) and inversely with ApoE3 genotype prevalence. Proteomics identified 120 differentially expressed proteins in anxious patients, enriched in oxidative phosphorylation and neurodegenerative pathways. CCDC126 emerged as a cadmium-associated biomarker, validated in rat models exposed to Cd. Combining CCDC126, blood Cd, Pb, and hypertension, a clinical prediction model achieved robust discrimination (AUC&#x2009;=&#x2009;0.80, validation cohort). CONCLUSIONS: This first integrative environmental-proteomic study highlights cadmium's synergistic role in anxiety pathophysiology and psychiatric comorbidity. The predictive model offers translatable potential for early risk stratification, while CCDC126 provides mechanistic insights for targeted interventions in populations exposed to environmental pollutants.

Cadmium↗

Survival prediction for clear cell renal cell carcinoma based on deep multimodal synergistic survival network.

Objective.To propose a deep multimodal synergistic survival analysis framework (Deep Multimodal Synergistic Survival Network, DMSSN) to achieve accurate prognostic analysis for clear cell renal cell carcinoma (ccRCC).Methods.This study (DMSSN) utilized matched multimodal data from the Cancer Genome Atlas-KIRC database, including CT imaging data, whole slide images, copy number variation (CNV) features, and clinical data. Deep Canonical Correlation Analysis was employed to map heterogeneous modalities into a shared latent space. Contrastive learning was introduced to enhance semantic consistency across multimodal features, and a gating network was utilized for the adaptive fusion of multimodal information to achieve precise survival risk prediction for patients.Results.Experimental results demonstrated that DMSSN achieved a Concordance Index (C-index) of 0.8153 &#xb1; 0.0994, with a Log-rank testp-value of 1.6553&#xd7;10-11. DMSSN exhibited significant performance advantages over traditional statistical methods like Log-rank-Cox (0.7055 &#xb1; 0.0670) and machine learning methods such as Random Survival Forest (RSF) (0.6836 &#xb1; 0.1048). Furthermore, in comparison with similar deep learning approaches, DMSSN outperformed late fusion strategies (0.7493 &#xb1; 0.1211) and discrete-time survival models such as DeepHit (0.7655 &#xb1; 0.1041) and Nnet-surv (0.7694 &#xb1; 0.0635). Notably, DMSSN still achieved the best predictive performance when compared to the classic deep survival model DeepSurv (0.7919 &#xb1; 0.0978) and advanced state-of-the-art multimodal fusion frameworks like Context-Aware Transformer (0.7735 &#xb1; 0.0818) and Multimodal Co-Attention Transformer (0.8102 &#xb1; 0.0972). Ablation studies showed that removing any single modality led to a decline in performance, with the largest numerical decrease occurring after removing CT imaging features (C-index decreased to 0.7327), validating the complementarity of multimodal data and the pivotal role of radiomic features in prognostic assessment. Module ablation experiments further confirmed the effectiveness of the core components.Conclusion:By effectively integrating imaging, pathology, genomic, and clinical features, the DMSSN framework demonstrates superior performance and robustness in the survival prediction of ccRCC.

Carcinoma, Renal Cell↗

Biological Parts in Yeast Synthetic Biology: From Regulatory Elements to Predictive Design Platforms.

Yeasts, particularly Saccharomyces cerevisiae, are important eukaryotic chassis for synthetic biology because of their tractable genetics, versatile toolkits, and broad utility in metabolic engineering and functional genomics. Progress in this field has been driven by biological parts that enable programmable control of gene expression and cellular behavior. Early efforts focused mainly on promoters, terminators, and other regulatory elements for tuning individual genes. However, as engineering expanded to multigene pathways, genetic circuits, and dynamic regulatory systems, the limits of part-centric design became clear. Part performance is often shaped by genomic context, chromatin state, host physiology, and interactions with other components, which restricts modularity and predictability. In response, yeast synthetic biology is shifting toward integrated design frameworks combining multilayer regulation, standardized assembly, automated experimentation, and computational modeling. This review provides an integrated perspective on the evolution of biological parts across DNA-, RNA-, and protein-level regulation, connecting these advances with assembly frameworks, biofoundries, and machine learning to trace the trajectory from part-centric engineering toward predictive, system-level design in yeast synthetic biology.

Biofoundry↗

Using classification tree and logistic regression methods to diagnose myocardial infarction.

Early and accurate diagnosis of myocardial infarction (MI) in patients who present to the Emergency Room (ER) complaining of chest pain is an important problem in emergency medicine. A number of decision aids have been developed to assist with this problem but have not achieved general use. Machine learning techniques, including classification tree and logistic regression (LR) methods, have the potential to create simple but accurate decision aids. Both a classification tree (FT Tree) and an LR model (FT LR) have been developed to predict the probability that a patient with chest pain is having an MI based solely upon data available at time of presentation to the ER. Training data came from a data set collected in Edinburgh, Scotland. Each model was then tested on a separate Edinburgh data set, as well as on a data set from a different hospital in Sheffield, England. Previously published models, the Goldman classification tree[1] and Kennedy LR equation[2], were evaluated on the same test data sets. On the Edinburgh test set, results showed that the FT Tree, FT LR, and Kennedy LR performed equally well, with ROC curve areas of 94.04%, 94.28%, and 94.30%, respectively, while the Goldman Tree's performance was significantly poorer, with an area of 84.03%. The difference in ROC areas between the first three models and the Goldman model is significant beyond the 0.0001 level. On the Sheffield test set, results showed that the FT Tree, FT LR, and Kennedy LR ROC areas were not significantly different (p > = 0.17), while the FT Tree again outperformed the Goldman Tree (p = 0.006). Unlike previous work[3], this study indicates that classification trees, which have certain advantages over LR models, may perform as well as LR models in the diagnosis of patients with MI.

Algorithms↗

Construction of a sequence motif characteristic of aminergic G protein-coupled receptors.

An approach to discover sequence patterns characteristic of ligand classes is described and applied to aminergic G protein-coupled receptors (GPCRs). Putative ligand-binding residue positions were inferred from considering three lines of evidence: conservation in the subfamily absent or underrepresented in the superfamily, any available mutation data, and the physicochemical properties of the ligand. For aminergic GPCRs, the motif is composed of a conserved aspartic acid in the third transmembrane (TM) domain (rhodopsin position 117) and a conserved tryptophan in the seventh TM domain (rhodopsin position 293); the roles of each are readily justified by molecular modeling of ligand-receptor interactions. This minimally defined motif is an appropriate computational tool for identifying additional, potentially novel aminergic GPCRs from a set of experimentally uncharacterized "orphan" GPCRs, complementing existing sequence matching, clustering, and machine-learning techniques. Motif sensitivity stems from the stepwise addition of residues characteristic of an entire class of ligand (and not tailored for any particular biogenic amine). This sensitivity is balanced by careful consideration of residues (evidence drawn from mutation data, correlation of ligand properties to residue properties, and location with respect to the extracellular face), thereby maintaining specificity for the aminergic class. A number of orphan GPCRs assigned to the aminergic class by this motif were later discovered to be a novel subfamily of trace amine GPCRs, as well as the successful classification of the histamine H4 receptor.

Amino Acid Sequence↗

Transcriptome Analysis, Machine Learning, and Experimental Identification of CDK7 Affecting the Progression of Pregnancy-induced Hypertension by Influencing Macrophage Polarization.

INTRODUCTION: Pregnancy-induced hypertension (PIH) is a severe pregnancy complication characterized by placental insufficiency, abnormal vascular remodeling, and immune dysregulation, but personalized therapeutic markers remain unclear. This study aimed to identify key genes and explore immune mechanisms in PIH using transcriptome analysis, machine learning, and experimental validation. METHODS: We analyzed the GSE204835 transcriptomic dataset to screen differentially expressed genes (DEGs) and performed Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and Gene Set Enrichment Analysis (GSEA) for functional annotation. Immune infiltration analysis was also performed to examine the immune landscape in PIH. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified key genes, which were validated in a PIH cell model. Flow cytometry and immunofluorescence assays assessed the effect of CDK7 knockdown on macrophage polarization. RESULTS: A total of 1,598 DEGs (1,123 upregulated, 475 downregulated) were identified. Enrichment analyses highlighted associations with embryonic organ development, oxidative phosphorylation, angiogenesis, and oxidative stress. Immune infiltration analysis revealed altered eosinophil and macrophage polarization in PIH. LASSO regression selected 12 key genes, with CDK7 showing the most significant upregulation in the PIH model. CDK7 knockdown promoted macrophage polarization toward the anti-inflammatory M2 phenotype. DISCUSSION: These findings link CDK7 to immune dysregulation in PIH by modulating macrophage polarization, expanding our understanding of PIH's molecular mechanisms. The study's limitations include reliance on public datasets and in vitro models, warranting in vivo validation. CONCLUSION: CDK7 emerges as a potential therapeutic target for PIH, offering new insights into immunoregulatory interventions for this complication.

Female↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Alzheimer's subtypes A supervised, unsupervised, multimodal, multilayered embedded recursive (SUMMER) AI study.

Since Alzheimer's disease (AD) is a heterogeneous disease, different subtypes may have distinct biological, genetic, and clinical characteristics, requiring tailored interventions. While several proposed subtypes of AD exist, there is still no clear consensus on a definitive classification. By leveraging complementary AI approaches, including supervised and unsupervised learning, within a recursive pipeline (SUMMER) that integrates multimodal datasets encompassing MRI measurements, phenotypes, and genetic data, our goal was to generate robust scientific evidence for identifying AD subtypes. Data was downloaded from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database and included neuroimaging data (MRI), genetics (SNPs), clinical diagnosis, and demographics. 1133 European American participants' images, aged 55-95, were included in this study. The analysis was multi-fold, where the first step involved applying an unsupervised application to a subset of the MRI sample (AD + cognitively normal (CN) aged matched groups, 100 men aged 68-85 years, and 76 women aged 68-85 years). The MRI brain gray matter was segmented into 44 regions of interest (ROIs) according to a standard atlas, and 618 features were extracted, including ROI voxel intensity measurements such as minimum, maximum, and histogram variables. Results identified a cluster of subtype AD men and a cluster of subtype AD women that were distinct from the rest of their respective samples. In the next step, the integrity of the identified subtype AD clusters was investigated using the XGBoost supervised machine learning application with genetic features (SNPs, N=36,724) and labels: the identified subtype AD cluster vs. the rest of the sample, stratified by sex. A significant AD subtype men model (accuracy=0.85, F1=0.72, AUC=0.83) and a significant women AD subtype model (accuracy=0.81, F1=0.81, AUC=0.81) were built, confirming the homogeneity of the isolated AD subtype clusters. Discriminative biomarkers were extracted from the significant models, including selected ROIs and SNPs. Finally, the subtype models were tested on an unseen subset of ADNI data. The genetic-based models identified clusters of AD subtype participants consisting of 34% of the men AD group and 47% of the women AD group. Phenotypic analysis indicates that lower body weight was associated with the women's AD subtype. Complex diseases like AD demand a sophisticated, multimodal approach for precise diagnosis. Effectively identifying disease subtypes enhances the potential for personalized treatment, ultimately improving patient outcomes.

Journal Article↗

Explaining the output of ensembles in medical decision support on a case by case basis.

The use of ensembles in machine learning (ML) has had a considerable impact in increasing the accuracy and stability of predictors. This increase in accuracy has come at the cost of comprehensibility as, by definition, an ensemble model is considerably more complex than its component models. This is of significance for decision support systems in medicine because of the reluctance to use models that are essentially black boxes. Work on making ensembles comprehensible has so far focused on global models that mirror the behaviour of the ensemble as closely as possible. With such global models there is a clear tradeoff between comprehensibility and fidelity. In this paper, we pursue another tack, looking at local comprehensibility where the output of the ensemble is explained on a case-by-case basis. We argue that this meets the requirements of medical decision support systems. The approach presented here identifies the ensemble members that best fit the case in question and presents the behaviour of these in explanation.

Anticoagulants↗

Large-scale predictions of secretory proteins from mammalian genomic and EST sequences.

Machine learning techniques have improved predictions of secretory proteins from protein, genomic and expressed sequence tag (EST) sequences. Artificial neural networks, physical sequence analysis using high-performance optimization, and hidden Markov models identify extremely variable signal peptides (the vehicles of protein transport across the endoplasmic reticulum membrane), transmembrane segments, and specific extracellular and intracellular domains as indicators of possible roles in the intercellular and intracellular chemical signaling pathways. The major role of peptide hormones, blood coagulation factors, carcinogenesis agents, and other secretory proteins in orchestrating multicellular life indicates pharmacological potential in the cure of major diseases and numerous biotechnological applications.

Animals↗

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient&#x2009;=&#x2009;-&#x2009;1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10&#xa0;years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung↗