Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A novel approach for short-term load forecasting using support vector machines.

A support vector machine (SVM) modeling approach for short-term load forecasting is proposed. The SVM learning scheme is applied to the power load data, forcing the network to learn the inherent internal temporal property of power load sequence. We also study the performance when other related input variables such as temperature and humidity are considered. The performance of our proposed SVM modeling approach has been tested and compared with feed-forward neural network and cosine radial basis function neural network approaches. Numerical results show that the SVM approach yields better generalization capability and lower prediction error compared to those neural network approaches.

Artificial Intelligence↗

Data-driven approaches in green microbiology: strategies for plant growth-promoting bacteria.

Plant growth-promoting bacteria (PGPB) are gaining attention as scalable biological solutions to enhance crop productivity and resilience. However, accurately identifying and characterizing PGPB remains challenging, particularly under variable environmental conditions where microbial functions are context-dependent and shaped by complex plant-microbe interactions. Advances in high-throughput sequencing have shifted the field from culture-dependent approaches to genome-informed strategies, enabling large-scale taxonomic and functional profiling. Although trait-based databases support the prediction of plant-beneficial genes, they capture only a fraction of the underlying biological complexity and often require labor-intensive analyses. Machine learning (ML) and deep learning (DL) have emerged as powerful tools to integrate genomic, physiological, and ecological data, enabling the prioritization of candidate strains with plant growth-promoting potential. To evaluate advances in the field, we conducted a systematic review of studies integrating ML and DL with PGPB characterization, assessing algorithm selection, performance, and target plant systems. Across 248 observations, only 6.0% of studies directly addressed PGPB screening, whereas the majority (77.4%) focused on plant disease detection, revealing a substantial gap in the application of AI to beneficial microorganisms for plant growth. Convolutional neural networks (CNNs) were the most frequently applied algorithms, largely driven by image-based phenotyping tasks. Overall, the field is constrained by limited datasets, high computational demands, and challenges in modeling multispecies and host-associated interactions. We highlight the need for integrative and interpretable ML and DL frameworks that bridge genomic data and functional validation. Such approaches represent a promising path toward scalable, data-driven discovery and deployment of bioinoculants in sustainable agriculture.

Agriculture↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-κB, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-κB, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Support vector analysis of color-Doppler images: a new approach for estimating indices of left ventricular function.

Reliable noninvasive estimators of global left ventricular (LV) chamber function remain unavailable. We have previously demonstrated a potential relationship between color-Doppler M-mode (CDMM) images and two basic indices of LV function: peak-systolic elastance (Emax) and the time-constant of LV relaxation (tau). Thus, we hypothesized that these two indices could be estimated noninvasively by adequate postprocessing of CDMM recordings. A semiparametric regression (SR) version of support vector machine (SVM) is here proposed for building a blind model, capable of analyzing CDMM images automatically, as well as complementary clinical information. Simultaneous invasive and Doppler tracings were obtained in nine mini-pigs in a high-fidelity experimental setup. The model was developed using a test and validation leave-one-out design. Reasonably acceptable prediction accuracy was obtained for both Emax (intraclass correlation coefficient Ric, = 0.81) and tau (Ric, = 0.61). For the first time, a quantitative, noninvasive estimation of cardiovascular indices is addressed by processing Doppler-echocardiography recordings using a learning-from-samples method.

Algorithms↗

An integrated machine learning system to computationally screen protein databases for protein binding peptide ligands.

A fairly large set of protein interactions is mediated by families of peptide binding domains, such as Src homology 2 (SH2), SH3, PDZ, major histocompatibility complex, etc. To identify their ligands by experimental screening is not only labor-intensive but almost futile in screening low abundance species due to the suppression by high abundance species. An ideal way of studying protein-protein interactions is to use high throughput computational approaches to screen protein sequence databases to direct the validating experiments toward the most promising peptides. Predictors with only good cross-validation were not good enough to screen protein databases. In the current study we built integrated machine learning systems using three novel coding methods and screened the Swiss-Prot and GenBank protein databases for potential ligands of 10 SH3 and three PDZ domains. A large fraction of predictions has already been experimentally confirmed by other independent research groups, indicating a satisfying generalization capability for future applications in identifying protein interactions.

Amino Acid Motifs↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Human pol II promoter prediction: time series descriptors and machine learning.

Although several in silico promoter prediction methods have been developed to date, they are still limited in predictive performance. The limitations are due to the challenge of selecting appropriate features of promoters that distinguish them from non-promoters and the generalization or predictive ability of the machine-learning algorithms. In this paper we attempt to define a novel approach by using unique descriptors and machine-learning methods for the recognition of eukaryotic polymerase II promoters. In this study, non-linear time series descriptors along with non-linear machine-learning algorithms, such as support vector machine (SVM), are used to discriminate between promoter and non-promoter regions. The basic idea here is to use descriptors that do not depend on the primary DNA sequence and provide a clear distinction between promoter and non-promoter regions. The classification model built on a set of 1000 promoter and 1500 non-promoter sequences, showed a 10-fold cross-validation accuracy of 87% and an independent test set had an accuracy >85% in both promoter and non-promoter identification. This approach correctly identified all 20 experimentally verified promoters of human chromosome 22. The high sensitivity and selectivity indicates that n-mer frequencies along with non-linear time series descriptors, such as Lyapunov component stability and Tsallis entropy, and supervised machine-learning methods, such as SVMs, can be useful in the identification of pol II promoters.

Algorithms↗

Bounds on error expectation for support vector machines.

We introduce the concept of span of support vectors (SV) and show that the generalization ability of support vector machines (SVM) depends on this new geometrical concept. We prove that the value of the span is always smaller (and can be much smaller) than the diameter of the smallest sphere containing the support vectors, used in previous bounds (Vapnik, 1998). We also demonstrate experimentally that the prediction of the test error given by the span is very accurate and has direct application in model selection (choice of the optimal parameters of the SVM).

Learning↗

Support vector machines for predicting the specificity of GalNAc-transferase.

Support Vector Machines (SVMs) which is one kind of learning machines, was applied to predict the specificity of GalNAc-transferase. The examination for the self-consistency and the jackknife test of the SVMs method were tested for the training dataset (305 oligopeptides), the correct rate of self-consistency and jackknife test reaches 100% and 84.9%, respectively. Furthermore, the prediction of the independent testing dataset (30 oligopeptides) was tested, the rate reaches 76.67%.

Algorithms↗

Discovery and performance of DNA methylation panels for cancer detection and classification in blood.

Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.

DNA methylation↗

Selection of patient samples and genes for outcome prediction.

Gene expression profiles with clinical outcome data enable monitoring of disease progression and prediction of patient survival at the molecular level. We present a new computational method for outcome prediction. Our idea is to use an informative subset of original training samples. This subset consists of only short-term survivors who died within a short period and long-term survivors who were still alive after a long follow-up time. These extreme training samples yield a clear platform to identify genes whose expression is related to survival. To find relevant genes, we combine two feature selection methods -- entropy measure and Wilcoxon rank sum test -- so that a set of sharp discriminating features are identified. The selected training samples and genes are then integrated by a support vector machine to build a prediction model, by which each validation sample is assigned a survival/relapse risk score for drawing Kaplan-Meier survival curves. We apply this method to two data sets: diffuse large-B-cell lymphoma (DLBCL) and primary lung adenocarcinoma. In both cases, patients in high and low risk groups stratified by our risk scores are clearly distinguishable. We also compare our risk scores to some clinical factors, such as International Prognostic Index score for DLBCL analysis and tumor stage information for lung adenocarcinoma. Our results indicate that gene expression profiles combined with carefully chosen learning algorithms can predict patient survival for certain diseases.

Biomarkers, Tumor↗

Evaluating variable selection methods for diagnosis of myocardial infarction.

This paper evaluates the variable selection performed by several machine-learning techniques on a myocardial infarction data set. The focus of this work is to determine which of 43 input variables are considered relevant for prediction of myocardial infarction. The algorithms investigated were logistic regression (with stepwise, forward, and backward selection), backpropagation for multilayer perceptrons (input relevance determination), Bayesian neural networks (automatic relevance determination), and rough sets. An independent method (self-organizing maps) was then used to evaluate and visualize the different subsets of predictor variables. Results show good agreement on some predictors, but also variability among different methods; only one variable was selected by all models.

Algorithms↗

Support-vector-machine classification of linear functional motifs in proteins.

Our algorithm predicts short linear functional motifs in proteins using only sequence information. Statistical models for short linear functional motifs in proteins are built using the database of short sequence fragments taken from proteins in the current release of the Swiss-Prot database. Those segments are confirmed by experiments to have single-residue post-translational modification. The sensitivities of the classification for various types of short linear motifs are in the range of 70%. The query protein sequence is dissected into short overlapping fragments. All segments are represented as vectors. Each vector is then classified by a machine learning algorithm (Support Vector Machine) as potentially modifiable or not. The resulting list of plausible post-translational sites in the query protein is returned to the user. We also present a study of the human protein kinase C family as a biological application of our method.

Databases, Genetic↗

A novel statistical ligand-binding site predictor: application to ATP-binding sites.

Structural genomics initiatives are leading to rapid growth in newly determined protein 3D structures, the functional characterization of which may still be inadequate. As an attempt to provide insights into the possible roles of the emerging proteins whose structures are available and/or to complement biochemical research, a variety of computational methods have been developed for the screening and prediction of ligand-binding sites in raw structural data, including statistical pattern classification techniques. In this paper, we report a novel statistical descriptor (the Oriented Shell Model) for protein ligand-binding sites, which utilizes the distance and angular position distribution of various structural and physicochemical features present in immediate proximity to the center of a binding site. Using the support vector machine (SVM) as the classifier, our model identified 69% of the ATP-binding sites in whole-protein scanning tests and in eukaryotic proteins the accuracy is particularly high. We propose that this feature extraction and machine learning procedure can screen out ligand-binding-capable protein candidates and can yield valuable biochemical information for individual proteins.

Adenosine Triphosphate↗

A robust meta-classification strategy for cancer diagnosis from gene expression data.

One of the major challenges in cancer diagnosis from microarray data is to develop robust classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose a meta-classification scheme which uses a robust multivariate gene selection procedure and integrates the results of several machine learning tools trained on raw and pattern data. We validate our method by applying it to distinguish diffuse large B-cell lymphoma (DLBCL) from follicular lymphoma (FL) on two independent datasets: the HuGeneFL Affmetrixy dataset of Shipp et al. (www. genome.wi.mit.du/MPR /lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our meta-classification technique achieves higher predictive accuracies than each of the individual classifiers trained on the same dataset and is robust against various data perturbations. We also find that combinations of p53 responsive genes (e.g., p53, PLK1 and CDK2) are highly predictive of the phenotype.

Algorithms↗

Maximum entropy modeling for mining patient medication status from free text.

Using a classification scheme of patient medication status we sought to recognize and categorize medications mentioned in the unrestricted text of clinical documents generated in clinical practice. The categories refer to the patient's status with respect to the medication such as discontinuation, start or initiation, and continuation of a given medication. This categorization is performed with a machine learning technique, Maximum Entropy (ME), that is well suited to incorporating heterogeneous sources of information necessary for classifying patient's medication status. We use hand labeled training data to generate ME models and test 5 different training feature sets. Our results show that the most optimal feature set includes a combination of the following: two words preceding and following the mention of the drug, the subject of the sentence in which the drug mention occurs, the 2 words following the subject, and a binary feature vector of lexicalized semantic cues indicative of medication status or its change. The average predictive power of a model trained on these features is approximately 89%.

Artificial Intelligence↗

Using feature generation and feature selection for accurate prediction of translation initiation sites.

Correct prediction of the translation initiation site (TIS) is an important issue in genomic research. We show that feature generation together with correlation based feature selection can be used with a variety of machine learning algorithms to give highly accurate translation initiation site prediction. Only very few features are needed and the results achieve comparable accuracy to the best existing approaches. Our approach has the advantage that it does not require one to devise a special prediction method; rather standard machine learning classifiers are shown to give very good performance on the selected features. The raw and generated features which we have found to be important are the following: positions -3 and -1 in the sequence; upstream k-grams for k=3, 4, and 5; stop-codon frequency; downstream in-frame 3-gram; and the distance of ATG to the beginning of the sequence. The best result, with an overall accuracy of 90%, is obtained by selecting only seven features from this set. The same features retrained with the use of a scanning model achieves an overall accuracy of 94% on this dataset.

Codon, Initiator↗

Protein fold class prediction: new methods of statistical classification.

Feed forward neural networks are compared with standard and new statistical classification procedures for the classification of proteins. We applied logistic regression, an additive model and projection pursuit regression from the methods based on a posterior probabilities; linear, quadratic and a flexible discriminant analysis from the methods based on class conditional probabilities, and the K-nearest-neighbors classification rule. Both, the apparent error rate obtained with the training sample (n = 143) and the test error rate obtained with the test sample (n = 125) and the 10-fold cross validation error were calculated. We conclude that some of the standard statistical methods are potent competitors to the more flexible tools of machine learning.

Algorithms↗