Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Feature mining and predictive model construction from severe trauma patient's data.

In management of severe trauma patients, trauma surgeons need to decide which patients are eligible for damage control. Such decision may be supported by utilizing models that predict the patient's outcome. The study described in this paper investigates the possibility to construct patient outcome prediction models from retrospective patient's data at the end of initial damage control surgery by using feature mining and machine learning techniques. As the data used comprises rather excessive number of features, special attention was paid to the problem of selecting only the most relevant features. We show that a small subset of features may carry enough information to construct reasonably accurate prognostic models. Furthermore, the techniques used in our study identified two factors, namely the pH value when admitted to ICU and the worst partial active thromboplastin time, to be of highest importance for prediction. This finding is pathophysiologically reasonable and represents two of three major problems with severe trauma patients, metabolic acidosis, hypothermia, and coagulopathy.

Algorithms↗

Analysis of respiratory pressure-volume curves in intensive care medicine using inductive machine learning.

We present a case study of machine learning and data mining in intensive care medicine. In the study, we compared different methods of measuring pressure-volume curves in artificially ventilated patients suffering from the adult respiratory distress syndrome (ARDS). Our aim was to show that inductive machine learning can be used to gain insights into differences and similarities among these methods. We defined two tasks: the first one was to recognize the measurement method producing a given pressure-volume curve. This was defined as the task of classifying pressure-volume curves (the classes being the measurement methods). The second was to model the curves themselves, that is, to predict the volume given the pressure, the measurement method and the patient data. Clearly, this can be defined as a regression task. For these two tasks, we applied C5.0 and CUBIST, two inductive machine learning tools, respectively. Apart from medical findings regarding the characteristics of the measurement methods, we found some evidence showing the value of an abstract representation for classifying curves: normalization and high-level descriptors from curve fitting played a crucial role in obtaining reasonably accurate models. Another useful feature of algorithms for inductive machine learning is the possibility of incorporating background knowledge. In our study, the incorporation of patient data helped to improve regression results dramatically, which might open the door for the individual respiratory treatment of patients in the future.

Adult↗

Instrumented Walkway Gait Analysis Predicts Fallers in Neurological Disorders: Identifying Digital Biomarkers for Balance Monitoring.

Assessing balance is crucial in neurological rehabilitation, yet while wearable sensors enable real-world monitoring, identifying reliable digital biomarkers remains challenging. This study utilized a high-fidelity instrumented walkway to determine which gait parameters best predict balance impairment, providing robust targets for future wearable applications. We analyzed 49 steady-state gait metrics from 140 individuals with diverse neurological conditions. Using statistical analysis and machine learning, we evaluated these parameters against objective force plate sway scores and clinical fall-history labels. Group analysis identified 16 parameters significantly distinguishing fallers from non-fallers, and a neural network classified fallers with an area under the curve of 0.75. Across all analytical approaches, overall gait variability, e.g., Stride Width S.D. and the Gait Variability Index, emerged as a universal predictor of balance impairment and fall risk. Furthermore, while traditional linear models emphasized spatial postural control, machine learning classification uniquely identified inter-limb asymmetry as a premier driver of fall prediction. These findings indicate that instrumented gait analysis effectively identifies digital biomarkers for balance deficits. Isolating these specific metrics provides a clear blueprint for meaningful metrics required for continuous objective monitoring and future development of personalized, adaptive rehabilitation strategies.

Humans↗

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping↗

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning↗

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data.

Long noncoding RNAs (lncRNAs) play critical regulatory roles in human development and disease. Although there are over 100,000 samples with available RNA sequencing (RNA-seq) data, many lncRNAs have yet to be annotated. The conventional approach to identifying novel lncRNAs from RNA-seq data is to find transcripts without coding potential but this approach has a false discovery rate of 30-75%. Other existing methods either identify only multi-exon lncRNAs, missing single-exon lncRNAs, or require transcriptional initiation profiling data (such as H3K4me3 ChIP-seq data), which is unavailable for many samples with RNA-seq data. Because of these limitations, current methods cannot accurately identify novel lncRNAs from existing RNA-seq data. To address this problem, we have developed software, Flnc, to accurately identify both novel and annotated full-length lncRNAs, including single-exon lncRNAs, directly from RNA-seq data without requiring transcriptional initiation profiles. Flnc integrates machine learning models built by incorporating four types of features: transcript length, promoter signature, multiple exons, and genomic location. Flnc achieves state-of-the-art prediction power with an AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the conventional approach to over 85%. Flnc is available via GitHub platform.

RNA-seq↗

Massively parallel approaches for characterizing noncoding functional variation in human evolution.

The genetic differences underlying unique phenotypes in humans compared to our closest primate relatives have long remained a mystery. Similarly, the genetic basis of adaptations between human groups during our expansion across the globe is poorly characterized. Uncovering the downstream phenotypic consequences of these genetic variants has been difficult, as a substantial portion lies in noncoding regions, such as cis-regulatory elements (CREs). Here, we review recent high-throughput approaches to measure the functions of CREs and the impact of variation within them. CRISPR screens can directly perturb CREs in the genome to understand downstream impacts on gene expression and phenotypes, while massively parallel reporter assays can decipher the regulatory impact of sequence variants. Machine learning has begun to be able to predict regulatory function from sequence alone, further scaling our ability to characterize genome function. Applying these tools across diverse phenotypes, model systems, and ancestries is beginning to revolutionize our understanding of noncoding variation underlying human evolution.

Humans↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape↗

Artificial Intelligence for Colorectal Surgeons-Part II: Research Applications, Challenges in Adoption, and Practical Resources.

BACKGROUND: This is part II of a 2-part series examining artificial intelligence in colorectal surgery. Part I established foundational concepts and clinical applications. Implementation, however, requires understanding research methodologies, available resources, and the specific challenges currently limiting widespread adoption. These topics are the focus of part II. OBJECTIVE: To examine artificial intelligence's transformation of surgical research, provide practical implementation resources, address adoption challenges, and explore future directions in colorectal surgery. METHODS: Comprehensive literature review focusing on artificial intelligence research methodology, implementation barriers, educational resources, and emerging technologies relevant to colorectal surgeons. RESULTS: Artificial intelligence streamlines clinical trial design through predictive modeling and natural language processing, reducing enrollment challenges that contribute to failed or inadequate trial accrual. Machine learning enables heterogeneity analysis within clinical trials, identifying treatment-responsive subgroups. Foundation models unlock analysis of unstructured electronic health record data at scale. Professional societies and universities offer specialized artificial intelligence education programs, with open-access data sets facilitating research participation. However, implementation faces multifaceted challenges: technical infrastructure demands, with real-time processing requiring dedicated graphics processing unit clusters; regulatory frameworks struggling with continuously evolving algorithms; undefined liability distribution for artificial intelligence-assisted decisions; algorithmic bias risking health care disparities; and the "black box" problem limiting clinical trust. Economic barriers include substantial initial costs without clear reimbursement pathways. Future directions include multimodal artificial intelligence integrating imaging, genomics, and histopathology; cognitive robotic systems with real-time decision support; digital twin technology for patient-specific surgical simulation; and global surgical artificial intelligence networks enabling distributed learning across institutions. CONCLUSIONS: Although artificial intelligence offers transformative potential for colorectal surgery research and practice, successful implementation requires addressing technical, regulatory, ethical, and economic challenges. The surgeon's evolving role demands both traditional expertise and computational fluency. Future advances in multimodal integration, autonomous systems, and global collaboration will fundamentally reshape surgical practice but will require thoughtful implementation prioritizing patient benefit and clinical value.

Humans↗

Donor Microbiota Features Associated With Liver Transplant Recipient Infectious Complications: A Pilot Study Using Deep Intestinal Sampling During Liver Procurement.

BACKGROUND: The gut microbiota of living organ donors has been linked to transplant outcomes. However, little is known about the characteristics of the deceased donor gut microbiota or its potential impact on recipient outcomes. METHODS: We analyzed the deep intestinal microbiota from 24 deceased donors. Samples included luminal stool from the right and left colon as well as bile. Microbial composition was characterized using 16S V4 rRNA sequencing. &#x3b1;- and &#x3b2;-diversity analyses were performed to compare microbial communities between donor enteric sites and against stool samples from 28 healthy community controls, 14 critically ill intensive care comparators, and 12 matched liver transplant recipients. Machine learning models and logistic regression analysis were applied to explore whether features of the donor microbiota could predict recipient post-transplant complications. FINDINGS: The deceased donor microbiota showed an absence of the expected compositional variability between sampling sites, with no significant differences in either &#x3b1;- or &#x3b2;-diversity observed between bile, right and left colonic samples (all p > 0.05). Donor samples exhibited distinct microbial profiles compared with stool from both healthy and ICU comparators, including increased abundance of potential pathogens within the Enterobacteriaceae family (all p < 0.001). Features of the donor microbiota, particularly enrichment of Enterobacteriaceae, were associated with an increased risk of early post-transplant infection in recipients (&#x2264;&#xa0;30 days; p&#xa0;=&#xa0;0.011). INTERPRETATION: The deceased donor gut microbiota may represent a distinct microbial community with potential clinical relevance. Microbial profiling of donor enteric microbiota may help identify recipients at heightened risk of early post-transplant infectious complications.

Enterobacteriaceae↗

Identification of Immune Response-Related Proteomic Biomarkers in Moyamoya Disease Using Serum Olink Proteomics.

Moyamoya disease, a rare chronic cerebrovascular disorder, requires invasive digital subtraction angiography (DSA) for diagnosis. This study employed high-throughput proteomics to identify plasma biomarkers for Moyamoya disease diagnosis. We conducted immunopanel analysis using the Olink platform to evaluate 92 immune-related proteins in plasma samples from 88 Moyamoya disease patients and 88 healthy controls. Key proteins were identified through differential expression analysis, GO, and KEGG enrichment analysis. A diagnostic model was constructed using LASSO regression, Boruta algorithm, and machine learning models including random forest and XGBoost. Validation of these proteins was performed using GEO external data sets, followed by prediction of potential therapeutic drugs and molecular docking validation through pharmacogenomic databases. A total of 44 differentially expressed proteins were identified through the Olink immunopanel, with 12 downregulated and 32 upregulated. GO and KEGG analyses revealed significant enrichment of these proteins in innate immune responses and signaling pathways such as NF-kB and MAPK. Through LASSO, random forest, and protein under-area analysis, four potential biomarkers for Moyamoya disease (MGMT, SIT1, PRDX1, TRAF2) were identified. A diagnostic model using these proteins showed the highest AUC value with the XGBoost model. Additionally, TRAF2 and PRDX1 exhibited significant expression differences in Moyamoya disease patients within the GEO data set. Our study revealed the immune landscape of Moyamoya disease, identified four biomarkers, and established a variety of diagnostic models.

Humans↗

On the use of machine learning to identify topological rules in the packing of beta-strands.

The machine learning program GOLEM was applied to discover topological rules in the packing of beta-sheets in alpha/beta-domain proteins. Rules (constraints) were determined for four features of beta-sheet packing: (i) whether a beta-strand is at an edge; (ii) whether two consecutive beta-strands pack parallel or anti-parallel; (iii) whether two beta-strands pack adjacently; and (iv) the winding direction of two consecutive beta-strands. Rules were found with high predictive accuracy and coverage. The errors were generally associated with complications in domain folds, especially in one doubly would domains. Investigation of the rules revealed interesting patterns, some of which were known previously, others that are novel. Novel features include (i) the relationship between pairs of sequential strands is in general one of decreasing size; (ii) more sequential pairs of strands wind in the direction out than in; and (iii) it takes a larger alteration in hydrophobicity to change a strand from winding in the direction out than in. These patterns in the data may be the result of folding pathways in the domains. The rules found are of predictive value and could be used in the combinatorial prediction of protein structure, or as a general test of model structures, e.g. those produced by threading. We conclude that machine learning has a useful role in the analysis of protein structures.

Amino Acid Sequence↗

Recent advances in computational prediction of drug absorption and permeability in drug discovery.

Approximately 40%-60% of developing drugs failed during the clinical trials because of ADME/Tox deficiencies. Virtual screening should not be restricted to optimize binding affinity and improve selectivity; and the pharmacokinetic properties should also be included as important filters in virtual screening. Here, the current development in theoretical models to predict drug absorption-related properties, such as intestinal absorption, Caco-2 permeability, and blood-brain partitioning are reviewed. The important physicochemical properties used in the prediction of drug absorption, and the relevance of predictive models in the evaluation of passive drug absorption are discussed. Recent developments in the prediction of drug absorption, especially with the application of new machine learning methods and newly developed software are also discussed. Future directions for research are outlined.

Computer Simulation↗

Distinguishing enzyme structures from non-enzymes without alignments.

The ability to predict protein function from structure is becoming increasingly important as the number of structures resolved is growing more rapidly than our capacity to study function. Current methods for predicting protein function are mostly reliant on identifying a similar protein of known function. For proteins that are highly dissimilar or are only similar to proteins also lacking functional annotations, these methods fail. Here, we show that protein function can be predicted as enzymatic or not without resorting to alignments. We describe 1178 high-resolution proteins in a structurally non-redundant subset of the Protein Data Bank using simple features such as secondary-structure content, amino acid propensities, surface properties and ligands. The subset is split into two functional groupings, enzymes and non-enzymes. We use the support vector machine-learning algorithm to develop models that are capable of assigning the protein class. Validation of the method shows that the function can be predicted to an accuracy of 77% using 52 features to describe each protein. An adaptive search of possible subsets of features produces a simplified model based on 36 features that predicts at an accuracy of 80%. We compare the method to sequence-based methods that also avoid calculating alignments and predict a recently released set of unrelated proteins. The most useful features for distinguishing enzymes from non-enzymes are secondary-structure content, amino acid frequencies, number of disulphide bonds and size of the largest cleft. This method is applicable to any structure as it does not require the identification of sequence or structural similarity to a protein of known function.

Algorithms↗

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans↗

Disease candidate genes prediction using positive labeled and unlabeled instances.

Identifying disease genes and understanding their performance is critical in producing drugs for genetic diseases. Nowadays, laboratory approaches are not only used for disease gene identification but also using computational approaches like machine learning are becoming considerable for this purpose. In machine learning methods, researchers can only use two data types (disease genes and unknown genes) to predict disease candidate genes. Notably, there is no source for the negative data set. The proposed method is a two-step process: The first step is the extraction of reliable negative genes from a set of unlabeled genes by one-class learning and a filter based on distance indicators from known disease genes; this step is performed separately for each disease. The second step is the learning of a binary model using causing genes of each disease as a positive learning set and the reliable negative genes extracted from that disease. Each gene in the unlabeled gene's production and ranking step is assigned a normalized score using two filters and a learned model. Consequently, disease genes are predicted and ranked. The proposed method evaluation of various six diseases and Cancer class indicates better results than other studies.

Humans↗

Mitochondria related gene signature serves as prognosis prediction and risk stratification of cholangiocarcinoma.

BACKGROUND: Cholangiocarcinoma (CHOL) is a highly aggressive biliary malignancy with poor clinical outcomes and limited effective prognostic biomarkers. Mitochondrial dysfunction participates in multiple oncological processes of CHOL, yet the prognostic roles of mitochondria&#x2011;related genes (MRGs) remain poorly understood. This study aimed to characterize MRGs expression in CHOL and develop a molecular prognostic model for predicting patient survival and guiding clinical management. METHODS: RNA sequencing (RNA-seq) and clinical data of CHOL were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) (GSE89748) databases. Differentially expressed MRGs were identified, and 10 machine learning algorithms were used to construct prognostic models. The optimal model (highest average C-index) was selected to establish a mitochondria-related risk score (MRRS), which was validated internally and externally. A nomogram integrating clinical factors and MRRS was developed, and biological mechanisms were explored via functional and immune analyses. RESULTS: A 3-MRG (MAP3K1, MRPL18, PYGB) prognostic signature was constructed, stratifying patients into high- and low-risk groups with significantly different overall survival. The model showed high predictive accuracy, with an area under the curve (AUC) up to 0.845, and MRRS was an independent prognostic factor. The signature was associated with mitochondrial pathways, and the high-risk group had distinct immune infiltration and mutation profiles. CONCLUSIONS: A validated MRG prognostic model effectively stratifies CHOL patients and has potential clinical value for prognosis prediction. Further validation in larger cohorts is needed to confirm its applicability.

Cholangiocarcinoma (CHOL)↗