Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Predicting the efficacy of short oligonucleotides in antisense and RNAi experiments with boosted genetic programming.

MOTIVATION: Both small interfering RNAs (siRNAs) and antisense oligonucleotides can selectively block gene expression. Although the two methods rely on different cellular mechanisms, these methods share the common property that not all oligonucleotides (oligos) are equally effective. That is, if mRNA target sites are picked at random, many of the antisense or siRNA oligos will not be effective. Algorithms that can reliably predict the efficacy of candidate oligos can greatly reduce the cost of knockdown experiments, but previous attempts to predict the efficacy of antisense oligos have had limited success. Machine learning has not previously been used to predict siRNA efficacy. RESULTS: We develop a genetic programming based prediction system that shows promising results on both antisense and siRNA efficacy prediction. We train and evaluate our system on a previously published database of antisense efficacies and our own database of siRNA efficacies collected from the literature. The best models gave an overall correlation between predicted and observed efficacy of 0.46 on both antisense and siRNA data. As a comparison, the best correlations of support vector machine classifiers trained on the same data were 0.40 and 0.30, respectively.

Algorithms↗

Genome-wide association, polygenic risk scores, and machine learning for chronic post-surgical pain risk stratification: A UK biobank study.

Chronic post-surgical pain is a prevalent and debilitating complication following surgery, representing a clinical challenge. Despite the established heritability of pain phenotypes, large-scale genetic studies remain limited. This study aimed to identify genetic variants associated with chronic post-surgical pain, develop polygenic risk scores, and integrate these with clinical features for risk prediction. UK Biobank data from 47,836 participants (2490 cases and 45,346 controls) were split into training (80%; n = 38,268) and validation (20%; n = 9568) sets prior to analysis. A genome-wide association study was conducted on the training set only, across 19 million variants, and polygenic risk scores were constructed and integrated with clinical features in a logistic regression framework. Two close, rare, imputed signals crossed the genome-wide significance threshold but lacked local linkage-disequilibrium support, while 220 variants crossed the suggestive threshold. In the held-out validation set, cases had higher mean polygenic risk scores than controls (0.138 vs. -0.021; Cohen's d = 0.16, p < 0.001). A logistic regression model integrating clinical features and polygenic risk scores achieved an area under the curve of 0.639 (95% CI: 0.583-0.693), higher than models using either feature set alone. The polygenic risk score for chronic post-surgical pain was among the most important predictors. Risk stratification revealed the top quartile had 3.84-fold higher odds of chronic post-surgical pain than the bottom quartile (95% CI: 2.00-7.37). These findings suggest a possible modest genetic contribution to chronic post-surgical pain. Polygenic risk scores may complement clinical factors in surgical risk stratification. PERSPECTIVE: Chronic post-surgical pain may have a modest genetic contribution. This UK Biobank study identified over 220 variants at suggestive significance and constructed a polygenic risk score that was significantly elevated in cases. A combined clinical-genomic model achieved a 3.84-fold difference in odds across predicted-risk quartiles.

Chronic post-surgical pain↗

A Risk Score for Polycystic Ovary Syndrome Based on Meta-Analysis and Machine Learning of Gut Microbiota Signatures.

Polycystic Ovary Syndrome (PCOS) is a prevalent endocrine and metabolic disorder among reproductive-age women, in which emerging evidence suggests a substantial role played by the gut microbiota. To comprehensively evaluate gut microbiota alterations in PCOS and identify microbial biomarkers through integrated analysis, a systematic search of PubMed, Web of Science, and Embase was conducted for studies employing 16S rRNA gene sequencing of fecal samples from PCOS cohorts. Ten eligible PCOS cohorts, comprising 858 individuals, were included in the study, from which a risk score was derived using a 20-gene gut microbial signature associated with PCOS. Meta-analysis at the genus level identified that Subdoligranulum, NK4A214_group, and Collinsella significantly decreased, and Bacteroides increased in PCOS across multiple cohorts. Machine learning analysis identified a 20-genus microbial signature using the least absolute shrinkage and selection operator (LASSO) method, which was used to construct a risk score with an AUC of 0.835 in diagnosis prediction. Network analysis further identified Negativibacillus and Lachnospiraceae_UCG_010 as potential driver microbes in PCOS. The analysis in this study highlights key alterations in the gut microbiota across PCOS cohorts. The identified gut microbial signature and derived LASSO-based risk model offer novel insights and a potential tool for PCOS diagnosis.

Polycystic Ovary Syndrome↗

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning↗

Advances in the prediction of protein targeting signals.

Enlarged sets of reference data and special machine learning approaches have improved the accuracy of the prediction of protein subcellular localization. Recent approaches report over 95% correct predictions with low fractions of false-positives for secretory proteins. A clear trend is to develop specifically tailored organism- and organelle-specific prediction tools rather than using one general method. Focus of the review is on machine learning systems, highlighting four concepts: the artificial neural feed-forward network, the self-organizing map (SOM), the Hidden-Markov-Model (HMM), and the support vector machine (SVM).

Animals↗

Using classification tree and logistic regression methods to diagnose myocardial infarction.

Early and accurate diagnosis of myocardial infarction (MI) in patients who present to the Emergency Room (ER) complaining of chest pain is an important problem in emergency medicine. A number of decision aids have been developed to assist with this problem but have not achieved general use. Machine learning techniques, including classification tree and logistic regression (LR) methods, have the potential to create simple but accurate decision aids. Both a classification tree (FT Tree) and an LR model (FT LR) have been developed to predict the probability that a patient with chest pain is having an MI based solely upon data available at time of presentation to the ER. Training data came from a data set collected in Edinburgh, Scotland. Each model was then tested on a separate Edinburgh data set, as well as on a data set from a different hospital in Sheffield, England. Previously published models, the Goldman classification tree[1] and Kennedy LR equation[2], were evaluated on the same test data sets. On the Edinburgh test set, results showed that the FT Tree, FT LR, and Kennedy LR performed equally well, with ROC curve areas of 94.04%, 94.28%, and 94.30%, respectively, while the Goldman Tree's performance was significantly poorer, with an area of 84.03%. The difference in ROC areas between the first three models and the Goldman model is significant beyond the 0.0001 level. On the Sheffield test set, results showed that the FT Tree, FT LR, and Kennedy LR ROC areas were not significantly different (p > = 0.17), while the FT Tree again outperformed the Goldman Tree (p = 0.006). Unlike previous work[3], this study indicates that classification trees, which have certain advantages over LR models, may perform as well as LR models in the diagnosis of patients with MI.

Algorithms↗

Predicting risk of coronary artery disease from DNA microarray-based genotyping using neural networks and other statistical analysis tool.

This paper presents a novel approach for complex disease prediction that we have developed, exemplified by a study on risk of coronary artery disease (CAD). This multi-disciplinary approach straddles fields of microarray technology and genetics, neural networks (NN), data mining and machine learning, as well as traditional statistical analysis techniques, namely principal components analysis (PCA) and factor analysis (FA). A description of the biological background of the study is given, followed by a detailed description of how the problem has been modeled for analyses by neural networks and FA. A committee learning approach for NN has been used to improve generalization rates. We show that our NN approach is able to yield promising prediction results despite using only the most fundamental network structures. More interestingly, through the statistical analysis process, genes of similar biological functions have been clustered. In addition, a gene marker involved in breaking down lipids has been found to be the most correlated to CAD.

Algorithms↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

ASGCL: Adaptive Sparse Mapping-based graph contrastive learning network for cancer drug response prediction.

Personalized cancer drug treatment is emerging as a frontier issue in modern medical research. Considering the genomic differences among cancer patients, determining the most effective drug treatment plan is a complex and crucial task. In response to these challenges, this study introduces the Adaptive Sparse Graph Contrastive Learning Network (ASGCL), an innovative approach to unraveling latent interactions in the complex context of cancer cell lines and drugs. The core of ASGCL is the GraphMorpher module, an innovative component that enhances the input graph structure via strategic node attribute masking and topological pruning. By contrasting the augmented graph with the original input, the model delineates distinct positive and negative sample sets at both node and graph levels. This dual-level contrastive approach significantly amplifies the model's discriminatory prowess in identifying nuanced drug responses. Leveraging a synergistic combination of supervised and contrastive loss, ASGCL accomplishes end-to-end learning of feature representations, substantially outperforming existing methodologies. Comprehensive ablation studies underscore the efficacy of each component, corroborating the model's robustness. Experimental evaluations further illuminate ASGCL's proficiency in predicting drug responses, offering a potent tool for guiding clinical decision-making in cancer therapy.

Humans↗

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article↗

Predicting protein-ligand binding affinities using novel geometrical descriptors and machine-learning methods.

Inspired by the concept of knowledge-based scoring functions, a new quantitative structure-activity relationship (QSAR) approach is introduced for scoring protein-ligand interactions. This approach considers that the strength of ligand binding is correlated with the nature of specific ligand/binding site atom pairs in a distance-dependent manner. In this technique, atom pair occurrence and distance-dependent atom pair features are used to generate an interaction score. Scoring and pattern recognition results obtained using Kernel PLS (partial least squares) modeling and a genetic algorithm-based feature selection method are discussed.

Algorithms↗

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a &#x3b2;-strand than an &#x3b1;-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans↗

In silico estimation of DMSO solubility of organic compounds for bioscreening.

Solubility of organic compounds in DMSO is an important issue for commercial and academic organizations handling large compound collections or performing biological screening. In particular, solubility data are critical for the optimization of storage conditions and for the selection of compounds for bioscreening compatible with the assay protocol. Solubility is largely determined by the solvation energy and the crystal disruption energy, and these molecular phenomena should be assessed in structure-solubility correlation studies. The authors summarize our long-term experimental observations and theoretical studies of physicochemical determinants of DMSO solubility of organic substances. They compiled a comprehensive reference database of proprietary data on compound solubility (55,277 compounds with good DMSO solubility and 10,223 compounds with poor DMSO solubility), calculated specific molecular descriptors (topological, electromagnetic, charge, and lipophilicity parameters), and applied an advanced machine-learning approach for training neural networks to address the solubility. Both supervised (feed-forward, back-propagated neural networks) and unsupervised (Kohonen neural networks) learning methods were used. The resulting neural network models were validated by successfully predicting DMSO solubility of compounds in independent test selections.

Dimethyl Sulfoxide↗

Natural language processing-based model to predict radiation pneumonitis in patients with locally advanced non-small cell lung cancer undergoing chemoradiotherapy: a retrospective cohort study.

BACKGROUND: Radiation pneumonitis (RP) remains a significant treatment-related toxicity in patients with unresectable, locally advanced non-small cell lung cancer (NSCLC) undergoing chemoradiotherapy (CRT). Most existing predictive models rely on static baseline demographic or dosimetry variables and lack real-time clinical applicability. We developed a novel predictive framework that integrates longitudinal symptom data extracted from clinical notes using natural language processing (NLP) with clinical and dosimetry features to improve early RP prediction. METHODS: We retrospectively identified 227 patients with locally advanced NSCLC treated with definitive CRT at a high-volume cancer center in the United States. We included all patients older than 18 years who were diagnosed between Jan 1, 2006, and Dec 31, 2022 with histologically or cytologically confirmed unresectable Stage 2 or 3 NSCLC and treated with conformal radiotherapy to a minimum dose of &#x2265;45 Gy with or without chemotherapy. Of these, 31 RP events were identified through manual adjudication using radiologic criteria and chart review. NLP was used to extract the temporal relationship of 16 pre-specified symptoms with treatment from over 100,000 clinical notes spanning pre- and during-treatment intervals. We trained and validated machine learning models on combinations of baseline clinical data, radiation dosimetry, and NLP-derived symptom features. Model performance was evaluated using a nested cross-validation framework, with an outer cross-validation loop reserved for performance assessment and an inner cross-validation loop used for model training and integration, and summarized using area under the receiver operating characteristic curve (AUC) and partial AUC (pAUC) at high specificity thresholds. Clinical utility was evaluated using decision curve analysis (DCA). FINDINGS: The best-performing model incorporated longitudinal NLP features and achieved a median AUC of 0.759 (90% confidence interval 0.753-0.766), significantly outperforming baseline models using only dosimetry (AUC 0.613) or clinical variables (AUC 0.635). NLP-based features such as cough trajectory, shortness of breath, and wheezing were among the most important predictors. Inclusion of NLP-derived symptom data improved early identification of high-risk patients, particularly in the clinically relevant high-specificity range (pAUC 0.021 vs. 0.010 for dosimetry alone). DCA showed that the calibrated MLP model provided greater net benefit than default strategies of treating all or no patients across clinically relevant threshold possibilities. INTERPRETATION: In this early work, NLP-based extraction of longitudinal symptoms from routine clinical documentation meaningfully enhances RP prediction in patients undergoing CRT for NSCLC. This approach leverages existing electronic health record infrastructure to deliver real-time, scalable, and interpretable risk estimates, offering a pathway toward potential early intervention and personalized toxicity management. The model and DCA requires external and prospective validation before clinical deployment; as such, future work should focus on this validation and integration into clinical decision support systems. FUNDING: AstraZeneca.

Chemoradiotherapy↗

Construction of a molecular diagnostic system for neurogenic rosacea by combining transcriptome sequencing and machine learning.

Patients with neurogenic rosacea (NR) frequently demonstrate pronounced neurological manifestations, often unresponsive to conventional therapeutic approaches. A molecular-level understanding and diagnosis of this patient cohort could significantly guide clinical interventions. In this study, we amalgamated our sequencing data (n&#x2009;=&#x2009;46) with a publicly accessible database (n&#x2009;=&#x2009;38) to perform an unsupervised cluster analysis of the integrated dataset. The eighty-four rosacea patients were partitioned into two distinct clusters. Neurovascular biomarkers were found to be elevated in cluster 1 compared to cluster 2. Pathways in cluster 1 were predominantly involved in neurotransmitter synthesis, transmission, and functionality, whereas cluster 2 pathways were centered on inflammation-related processes. Differential gene expression analysis and WGCNA were employed to delineate the characteristic gene sets of the two clusters. Subsequently, a diagnostic model was constructed from the identified gene sets using linear regression methodologies. The model's C index, comprising genes PNPLA3, CUX2, PLIN2, and HMGCR, achieved a remarkable value of 0.9683, with an area under the curve (AUC) for the training cohort's nomogram of 0.9376. Clinical characteristics from our dataset (n&#x2009;=&#x2009;46) were assessed by three seasoned dermatologists, forming the NR validation cohort (NR, n&#x2009;=&#x2009;18; non-neurogenic rosacea, n&#x2009;=&#x2009;28). Upon application of our model to NR diagnosis, the model's AUC value reached 0.9023. Finally, potential therapeutic candidates for both patient groups were predicted via the Connectivity Map. In summation, this study unveiled two clusters with unique molecular phenotypes within rosacea, leading to the development of a precise diagnostic model instrumental in NR diagnosis.

Humans↗

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-&#x3b1;/NF-&#x3ba;B, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans↗

Linear regression models for solvent accessibility prediction in proteins.

The relative solvent accessibility (RSA) of an amino acid residue in a protein structure is a real number that represents the solvent exposed surface area of this residue in relative terms. The problem of predicting the RSA from the primary amino acid sequence can therefore be cast as a regression problem. Nevertheless, RSA prediction has so far typically been cast as a classification problem. Consequently, various machine learning techniques have been used within the classification framework to predict whether a given amino acid exceeds some (arbitrary) RSA threshold and would thus be predicted to be "exposed," as opposed to "buried." We have recently developed novel methods for RSA prediction using nonlinear regression techniques which provide accurate estimates of the real-valued RSA and outperform classification-based approaches with respect to commonly used two-class projections. However, while their performance seems to provide a significant improvement over previously published approaches, these Neural Network (NN) based methods are computationally expensive to train and involve several thousand parameters. In this work, we develop alternative regression models for RSA prediction which are computationally much less expensive, involve orders-of-magnitude fewer parameters, and are still competitive in terms of prediction quality. In particular, we investigate several regression models for RSA prediction using linear L1-support vector regression (SVR) approaches as well as standard linear least squares (LS) regression. Using rigorously derived validation sets of protein structures and extensive cross-validation analysis, we compare the performance of the SVR with that of LS regression and NN-based methods. In particular, we show that the flexibility of the SVR (as encoded by metaparameters such as the error insensitivity and the error penalization terms) can be very beneficial to optimize the prediction accuracy for buried residues. We conclude that the simple and computationally much more efficient linear SVR performs comparably to nonlinear models and thus can be used in order to facilitate further attempts to design more accurate RSA prediction methods, with applications to fold recognition and de novo protein structure prediction methods.

Amino Acids↗

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans↗