Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Predicting gene function in Saccharomyces cerevisiae.

MOTIVATION: S.cerevisiae is one of the most important model organisms, and has has been the focus of over a century of study. In spite of these efforts, 40% of its open reading frames (ORFs) remain classified as having unknown function (MIPS: Munich Information Center for Protein Sequences). We wished to make predictions for the function of these ORFs using data mining, as we have previously successfully done for the genomes of M.tuberculosis and E.coli. Applying this approach to the larger and eukaryotic S.cerevisiae genome involves modifying the machine learning and data mining algorithms, as this is a larger organism with more data available, and a more challenging functional classification. RESULTS: Novel extensions to the machine learning and data mining algorithms have been devised in order to deal with the challenges. Accurate rules have been learned and predictions have been made for many of the ORFs whose function is currently unknown. The rules are informative, agree with known biology and allow for scientific discovery. AVAILABILITY: All predictions are freely available from http://www.genepredictions.org, all datasets used in this study are freely available from http://www.aber.ac.uk/compsci/Research/bio/dss/yeastdataand software for relational data mining is available from http://www.aber.ac.uk/compsci/Research/bio/dss/polyfarm.

Chromosome Mapping↗

Conserved HSFA1-dependent chromatin dynamics drive heat stress responses in plants.

Eukaryotic organisms remodel chromatin landscapes to regulate gene expression in response to environmental stress. In plants, heat stress (HS) induces widespread chromatin changes, yet the role of heat shock transcription factors (HSFs) in chromatin remodeling and their evolutionary conservation remains unclear. Using Marchantia polymorpha Mphsf mutants and Arabidopsis thaliana Athsfa1s mutants, we identify HSFA1 as a key regulator of HS-induced cis-regulatory element (CRE) accessibility, a mechanism conserved across land plants, mice, and humans. Gene regulatory network modeling reveals parallel transcription factor subnetworks, with MpWRKY10 and MpABI5B acting as indirect and negative HS regulators. We further showed that ABA modulates gene expression in an HSFA1-dependent manner without inducing chromatin remodeling. Finally, we develop a machine learning framework integrating chromatin accessibility and CRE information to predict gene expression across species, revealing stress-responsive regulatory logic at the transcriptional level. These findings provide insights into how TFs coordinate chromatin architecture to drive stress adaptation.

Heat-Shock Response↗

Crosstalk between cysteine and lysine modifications: Integrating redox and metabolic regulation.

Protein post-translational modifications (PTMs) on amino acid residues enable dynamic cellular responses to changes in metabolic and redox state. Cysteine and lysine are among the most extensively modified amino acid residues, with both undergoing a diversity of acylation and oxidative modifications. Indeed, proximal (<10&#x202f;&#xc5;) cysteine and lysine residues may form integration nodes for crosstalk between metabolism and redox homeostasis pathways. This review highlights the interaction of proximal Cys-Lys residues, including influence on residue pKa by local electrostatics, cysteine-to-lysine transfer of PTM moieties, and covalent crosslinking. We discuss candidate Cys-Lys regulatory pairs in proteins involved in redox regulation, proteostasis, metabolic adaptation and inflammation. We further utilize computational modeling to identify proximity between cysteine and lysine residues in proteins known to be regulated by acylation and oxidative PTMs, and to demonstrate changes in these distances and local electrostatic potential due to lysine acetylation. Finally, we review how mass spectrometry-based proteomics and machine-learning PTM predictive tools can enable the identification, validation, and interpretation of proximal Cys-Lys interactions that regulate cellular responses to oxidative challenge and metabolic flux.

Cysteine↗

Data mining in bioinformatics using Weka.

UNLABELLED: The Weka machine learning workbench provides a general-purpose environment for automatic classification, regression, clustering and feature selection-common data mining problems in bioinformatics research. It contains an extensive collection of machine learning algorithms and data pre-processing methods complemented by graphical user interfaces for data exploration and the experimental comparison of different machine learning techniques on the same problem. Weka can process data given in the form of a single relational table. Its main objectives are to (a) assist users in extracting useful information from data and (b) enable them to easily identify a suitable algorithm for generating an accurate predictive model from it. AVAILABILITY: http://www.cs.waikato.ac.nz/ml/weka.

Algorithms↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

RAS signaling in lung adenocarcinoma is defined by lineage context and DUSP4 loss.

BACKGROUNDThe molecular landscape of lung adenocarcinoma (LUAD) is often illustrated as a driver-oncogene pie chart, but identical mutations exhibit heterogeneous signaling shaped by comutations, transcriptional programs, and lineage context. We propose a lineage-integrated signaling framework using an EGFR mutation signature (mSig).METHODSWe defined EGFR mSig using differentially expressed genes in EGFR-mutant (EGFR-mt) LUADs. Semisupervised clustering and machine learning models were used to test reproducibility in different combinations of datasets. We analyzed molecular subtypes, lineage markers, co-occurring mutations, and EGFR copy number alterations in EGFR mSig-defined subtypes of LUAD.RESULTSEGFR mSig showed robust classification performance (area under receiver operating characteristic curve = 0.83-0.95; mean negative predictive value = 96.3%). Validated gene expression subtypes and lung lineage markers were closely aligned with EGFR mSig status. Most EGFR mSig+ tumors, including many without EGFR mutations, belonged to the bronchioid subtype. A subset of canonical RAS mutations were mSig+ and mirrored the EGFR mutation pattern. EGFR WT/mSig- tumors were enriched for nonbronchioid subtypes and had comutations in TP53 or RAS/RAF/RTKs. We highlight a parsimonious collection of coordinated mutations, including RAS, KEAP1, STK11, TP53, and CDKN2A, that taken together suggest coordination of tumor signaling previously suggested but now reproduced and expanded.CONCLUSIONA potentially novel EGFR mSig that captures the transcriptional footprint of EGFR activation revealed a subset of EGFR WT LUADs with mt-like features. mSig refines LUAD taxonomy beyond mutation-only pie-chart models by incorporating lineage and comutation context. Lineage-directed stratification with coalteration identifies clinically relevant groups across EGFR and RAS states and highlights treatment opportunities for patients currently considered oncogene-negative.FUNDINGNational Cancer Institute (NCI) U01CA272541, R01CA262296, U24CA264021, UG1CA233333, R01CA211939.

Humans↗

Genomic data sampling and its effect on classification performance assessment.

BACKGROUND: Supervised classification is fundamental in bioinformatics. Machine learning models, such as neural networks, have been applied to discover genes and expression patterns. This process is achieved by implementing training and test phases. In the training phase, a set of cases and their respective labels are used to build a classifier. During testing, the classifier is used to predict new cases. One approach to assessing its predictive quality is to estimate its accuracy during the test phase. Key limitations appear when dealing with small-data samples. This paper investigates the effect of data sampling techniques on the assessment of neural network classifiers. RESULTS: Three data sampling techniques were studied: Cross-validation, leave-one-out, and bootstrap. These methods are designed to reduce the bias and variance of small-sample estimations. Two prediction problems based on small-sample sets were considered: Classification of microarray data originating from a leukemia study and from small, round blue-cell tumours. A third problem, the prediction of splice-junctions, was analysed to perform comparisons. Different accuracy estimations were produced for each problem. The variations are accentuated in the small-data samples. The quality of the estimates depends on the number of train-test experiments and the amount of data used for training the networks. CONCLUSION: The predictive quality assessment of biomolecular data classifiers depends on the data size, sampling techniques and the number of train-test experiments. Conservative and optimistic accuracy estimations can be obtained by applying different methods. Guidelines are suggested to select a sampling technique according to the complexity of the prediction problem under consideration.

Computational Biology↗

AI-driven CRISPR screening: optimizing gene editing through automation and intelligent decision support.

BACKGROUND: CRISPR-based genetic screening has become a central methodology in functional genomics, enabling systematic interrogation of gene function, genetic interactions and context-dependent vulnerabilities at scale. However, the rapid expansion of screening modalities-including multi-condition designs, combinatorial perturbations, in vivo applications and single-cell readouts-has exposed fundamental limitations of heuristic-driven experimental design and post hoc statistical analysis. MAIN BODY: This Review synthesizes how artificial intelligence is reshaping CRISPR screening by introducing predictive, adaptive and system-level intelligence across the experimental lifecycle. We organize recent advances into two tightly coupled modules. First, machine learning and deep learning (ML/DL) methods optimize experimental design by learning context-dependent perturbation behavior, anticipating confounding effects and enabling iterative, information-efficient screening strategies. Second, large language model-agent (LLM-agent) systems complement these advances by externalizing scientific reasoning, integrating biological knowledge at scale and coordinating analysis and decision-making in human-in-the-loop workflows. CONCLUSIONS: Together, ML/DL and LLM-agent approaches reframe CRISPR screening from a static analytical pipeline into an intelligent experimental system, with important implications for robustness, scalability and biological discovery.

Artificial Intelligence↗

Application of causal discovery of factors driving dissolved oxygen in estuarine environments.

Dissolved oxygen (DO) concentrations in estuarine bottom waters are a manifestation of multiple, interacting physical and biogeochemical processes, yet identifying their independent contributions remains challenging. Here, we analyze monthly water quality monitoring data from eight stations across Long Island Sound from 1994 to 2022 using a causal discovery framework (PCMCI+) and transformation of forcing variables. Our goal is to identify and isolate variables that causally influence bottom DO and improve predictive models by minimizing overfitting and multicollinearity. PCMCI+ reveals surface-layer temperature as the most important and consistent negative driver of bottom DO, followed by stratification. Wind events exhibit only brief relief by advection and mixing, while river discharge shows no direct causal link to DO, making it less influential than previously thought. Biogeochemical variables, including chlorophyll-a (Chl-a), nitrate and nitrite, and particulate carbon, influence DO through both contemporaneous and time-lagged pathways, often with signs that shift depending on the process. The derived models were evaluated by comparing skill scores, mean squared error, and Akaike Information Criterion. Both model types perform well, with coefficient of determination values exceeding 0.90 at multiple stations using only 3-5 predictors. Our analysis reveals that the best causal predictors are surface-layer temperature, stratification, Chl-a, and particle carbon. This approach provides a scalable framework for improving prediction models and understanding the mechanistic links that control the seasonal variability of DO in estuarine systems.

Estuaries↗

APNet, an explainable sparse deep learning model to discover differentially active drivers of severe COVID-19.

MOTIVATION: Computational analyses of bulk and single-cell omics provide translational insights into complex diseases, such as COVID-19, by revealing molecules, cellular phenotypes, and signalling patterns that contribute to unfavourable clinical outcomes. Current in silico approaches dovetail differential abundance, biostatistics, and machine learning, but often overlook nonlinear proteomic dynamics, like post-translational modifications, and provide limited biological interpretability beyond feature ranking. RESULTS: We introduce APNet, a novel computational pipeline that combines differential activity analysis based on SJARACNe co-expression networks with PASNet, a biologically informed sparse deep learning model, to perform explainable predictions for COVID-19 severity. The APNet driver-pathway network ingests SJARACNe co-regulation and classification weights to aid result interpretation and hypothesis generation. APNet outperforms alternative models in patient classification across three COVID-19 proteomic datasets, identifying predictive drivers and pathways, including some confirmed in single-cell omics and highlighting under-explored biomarker circuitries in COVID-19. AVAILABILITY AND IMPLEMENTATION: APNet's R, Python scripts, and Cytoscape methodologies are available at https://github.com/BiodataAnalysisGroup/APNet.

COVID-19↗

Machine learning approaches to lung cancer prediction from mass spectra.

We addressed the problem of discriminating between 24 diseased and 17 healthy specimens on the basis of protein mass spectra. To prepare the data, we performed mass to charge ratio (m/z) normalization, baseline elimination, and conversion of absolute peak height measures to height ratios. After preprocessing, the major difficulty encountered was the extremely large number of variables (1676 m/z values) versus the number of examples (41). Dimensionality reduction was treated as an integral part of the classification process; variable selection was coupled with model construction in a single ten-fold cross-validation loop. We explored different experimental setups involving two peak height representations, two variable selection methods, and six induction algorithms, all on both the original 1676-mass data set and on a prescreened 124-mass data set. Highest predictive accuracies (1-2 off-sample misclassifications) were achieved by a multilayer perceptron and Naïve Bayes, with the latter displaying more consistent performance (hence greater reliability) over varying experimental conditions. We attempted to identify the most discriminant peaks (proteins) on the basis of scores assigned by the two variable selection methods and by neural network based sensitivity analysis. These three scoring schemes consistently ranked four peaks as the most relevant discriminators: 11683, 1403, 17350 and 66107.

Algorithms↗

Predicting essential genes in fungal genomes.

Essential genes are required for an organism's viability, and the ability to identify these genes in pathogens is crucial to directed drug development. Predicting essential genes through computational methods is appealing because it circumvents expensive and difficult experimental screens. Most such prediction is based on homology mapping to experimentally verified essential genes in model organisms. We present here a different approach, one that relies exclusively on sequence features of a gene to estimate essentiality and offers a promising way to identify essential genes in unstudied or uncultured organisms. We identified 14 characteristic sequence features potentially associated with essentiality, such as localization signals, codon adaptation, GC content, and overall hydrophobicity. Using the well-characterized baker's yeast Saccharomyces cerevisiae, we employed a simple Bayesian framework to measure the correlation of each of these features with essentiality. We then employed the 14 features to learn the parameters of a machine learning classifier capable of predicting essential genes. We trained our classifier on known essential genes in S. cerevisiae and applied it to the closely related and relatively unstudied yeast Saccharomyces mikatae. We assessed predictive success in two ways: First, we compared all of our predictions with those generated by homology mapping between these two species. Second, we verified a subset of our predictions with eight in vivo knockouts in S. mikatae, and we present here the first experimentally confirmed essential genes in this species.

Computational Biology↗

What can we learn from noncoding regions of similarity between genomes?

BACKGROUND: In addition to known protein-coding genes, large amounts of apparently non-coding sequence are conserved between the human and mouse genomes. It seems reasonable to assume that these conserved regions are more likely to contain functional elements than less-conserved portions of the genome. METHODS: Here we used a motif-oriented machine learning method based on the Relevance Vector Machine algorithm to extract the strongest signal from a set of non-coding conserved sequences. RESULTS: We successfully fitted models to reflect the non-coding sequences, and showed that the results were quite consistent for repeated training runs. Using the learned models to scan genomic sequence, we found that they often made predictions close to the start of annotated genes. We compared this method with other published promoter-prediction systems, and showed that the set of promoters which are detected by this method is substantially similar to that detected by existing methods. CONCLUSIONS: The results presented here indicate that the promoter signal is the strongest single motif-based signal in the non-coding functional fraction of the genome. They also lend support to the belief that there exists a substantial subset of promoter regions which share several common features including, but not restricted to, a relative abundance of CpG dinucleotides. This subset is detectable by a variety of distinct computational methods.

Animals↗

Machine learning approaches for phenotype-genotype mapping: predicting heterozygous mutations in the CYP21B gene from steroid profiles.

OBJECTIVE: Non-linear relations between multiple biochemical parameters are the basis for the diagnosis of many diseases. Traditional linear analytical methods are not reliable predictors. Novel nonlinear techniques are increasingly used to improve the diagnostic accuracy of automated data interpretation. This has been exemplified in particular for the classification and diagnostic prediction of cancers based on expression profiling data. Our objective was to predict the genotype from complex biochemical data by comparing the performance of experienced clinicians to traditional linear analysis, and to novel non-linear analytical methods. DESIGN AND METHODS: As a model, we used a well-defined set of interconnected data consisting of unstimulated serum levels of steroid intermediates assessed in 54 subjects heterozygous for a mutation of the 21-hydroxylase gene (CYP21B) and in 43 healthy controls. RESULTS: The genetic alteration was predicted from the pattern of steroid levels with an accuracy of 39% by clinicians and of 64% by linear analysis. In contrast, non-linear analysis, such as self-organizing artificial neural networks, support vector machines, and nearest neighbour classifiers, allowed for higher accuracy up to 83%. CONCLUSIONS: The successful application of these non-linear adaptive methods to capture specific biochemical problems may have generalized implications for biochemical testing in many areas. Nonlinear analytical techniques such as neural networks, support vector machines, and nearest neighbour classifiers may serve as an important adjunct to the decision process of a human investigator not 'trained' in a specific complex clinical or laboratory setting and may aid them to classify the problem more directly.

Adult↗

Demonstrating the potential of untargeted hair proteomics for personalized biomarkers in stress-associated disorders.

Biomarker research in psychopathology increasingly employs high-dimensional Omics approaches. Yet, proteomics based on human hair remain largely unexplored, despite its potential to efficiently capture stable biological signals accumulated over weeks to months. This study leveraged machine learning to investigate the potential of the hair proteome-all detectable peptides and proteins-as a biomarker source for stress-associated psychopathology. We analyzed protein profiles from hair segments of women with non-suicidal self-injury disorder and healthy controls (N&#x202f;=&#x202f;68). Of 1114 identified proteins, 611 were sufficiently abundant for analyses. Partial Least Squares Discriminant Analysis achieved stable 84.4&#xa0;% cross-validated accuracy for classification of clinical groups (p&#x202f;<&#x202f;.001), outperforming models based on data-derived clusters (60&#xa0;%), stress-related proteins (73&#xa0;%), and simulated hair cortisol from meta-analytic effect sizes (53-59&#xa0;%). Predicted class probabilities strongly correlated with clinical symptoms and well-being (r&#x202f;>&#x202f;.60). Key predictive proteins were linked to pain perception, oxidative stress, and cholesterol homeostasis. Approximately 15&#xa0;% of proteins differed significantly between groups, with the strongest candidates related to ribosomal function-an emerging target in depression. These findings establish hair proteomics as a promising, non-invasive biomarker source for psychiatric research with potential clinical applications in risk assessment and personalized interventions.

Humans↗

Predicting the efficiency of UAG translational stop signal through studies of physicochemical properties of its composite mono- and dinucleotides.

In this study, we explored the problem of predicting the UAG stop-codon read-through efficiency. The reported nucleotide sequences were first converted into physicochemical property vectors before being presented to a machine learning algorithm. Two sets of physicochemical properties were applied: one for mononucleosides (in terms of steric bulk, hydrophobicity and electronics) and another for dinucleotides. To the best of our knowledge, this is the first report of how dinucleotides are converted into principle components derived from NMR chemical shift data. A few efficiency prediction models were then derived and a comparison between mononucleoside and dinucleotide-based models was shown. In the derived models, the coefficients of these property based predictors lend themselves to bio-physical interpretations, an advantage which is demonstrated in this study via a prediction model based on the steric bulk factor. Although it is quite simple, the steric bulk factor model explained well the effect of sequence variations surrounding the amber stop codon and the tRNA bearing UCCU anticodon. We further proposed new alternatives at position -1 and +4 of a UAG stop codon sequence to enhance the readthrough efficiency. This research may contribute to a better understanding of the readthrough mechanisms and may also help to study the normal translation termination process.

Codon, Terminator↗

Quantitative structure-toxicity relationships (QSTRs): a comparative study of various non linear methods. General regression neural network, radial basis function neural network and support vector machine in predicting toxicity of nitro- and cyano- aromatics to Tetrahymena pyriformis.

Prediction of toxicity of 203 nitro- and cyano-aromatic chemicals to Tetrahymena pyriformis was carried out by radial basis function neural network, general regression neural network and support vector machine, in non-linear response surface methodology. Toxicity was predicted from hydrophobicity parameter (log Kow) and maximum superdelocalizability (Amax). Special attention was drawn to prediction ability and robustness of the models, investigated both in a leave-one-out and 10-fold cross validation (CV) processes. The influence that the corresponding changes in the learning sets during these CV processes could have on a common external test set including 41 compounds was also examined. This allowed us to establish the stability of the models. The non linear results slightly outperform (as expected) multilinear relationships (MLR) and also favourably compete with various other non linear approaches recently proposed by Ren (J. Chem. Inf. Comput. Sci., 43 1679 (2003)).

Animals↗

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals↗