Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Exploring the mechanism of aroma production in fermented cherry juice by L. brevis LD1.0600 using flavomics and whole genome analysis.

This study focused on L.brevis LD1.0600 with excellent fermentation traits: it analyzed genome-wide key regulatory genes for micro-metabolites, combined with fermented cherry juice flavor metabolomics data, and used machine learning to explore correlations between gene regulation, metabolite production, and flavor formation. The SVM model screened and verified fermented cherry juice VOCs; through OAV and flavor wheel analysis, LD1.0600 emerged as the top-performing strain, with a sweet, fruity dominant aroma. Key aroma-active components (OAV > 100) included 2-methoxy-4-vinylphenol, benzaldehyde, 2-methyl-butanoic acid and hexanoic acid, and 2-methoxy-4-vinylphenol and hexanoic acid elevated by LD1.0600-regulated genes (Chrom1-001884, Chrom1-000925, fabF and Chrom1-000199). At the same time, through research, a "strain screening-SVM screening of DVCs-OAV screening of key aroma components-whole genome sequencing of flavor regulatory genes" system was established. This system can not only be applied to the screen fermentation strains, but also can be extended to the application of other fermentation products.

Fermentation↗

Metric learning for text documents.

Many algorithms in machine learning rely on being given a good distance metric over the input space. Rather than using a default metric such as the Euclidean metric, it is desirable to obtain a metric based on the provided data. We consider the problem of learning a Riemannian metric associated with a given differentiable manifold and a set of points. Our approach to the problem involves choosing a metric from a parametric family that is based on maximizing the inverse volume of a given data set of points. From a statistical perspective, it is related to maximum likelihood under a model that assigns probabilities inversely proportional to the Riemannian volume element. We discuss in detail learning a metric on the multinomial simplex where the metric candidates are pull-back metrics of the Fisher information under a Lie group of transformations. When applied to text document classification the resulting geodesic distance resemble, but outperform, the tfidf cosine similarity measure.

Algorithms↗

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms↗

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals↗

DNA Methylation-Based Classification of Kidney Neoplasms.

Renal neoplasms are morphologically and molecularly heterogeneous, with their diagnosis often hindered by interobserver variability and overlapping microscopic features. A subset of cases is unclassifiable despite immunohistochemical, mutation, and cytogenetic-based diagnostic workup. Through examination of the genome-wide DNA methylation signatures of over 2000 renal neoplasms, we identified 23 coherent groups that correlate with known neoplasm types and identified novel clinically relevant subtypes of existing neoplasm types. We used machine learning models to develop and validate a classifier trained on DNA methylation profiles of 1284 samples. The classifier was tested on an external data set of 287 renal neoplasms with >90% concordance between expected neoplasm type and high-score DNA methylation-based classification. Discordance between the original histologic label and methylation class led to potential reclassification of some cases. This work demonstrates proof of principle for the feasibility of a DNA methylation classifier as a clinically useful tool to assist in the diagnosis of renal neoplasms.

Humans↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

Minimizing Off-Target Effects of CRISPR-Cas9 With Optimized sgRNA: Evaluation of Efficiency and Specificity in the Tumor Protein 53 (TP53) Region.

CRISPR-Cas9 is a widely used genetic tool with therapeutic potential in molecular biology. CRISPR-Cas9 enables precise genome editing by its ability to target specific DNA sequence. After off-target and on-target regions are identified, CRISPR-Cas9 is applied to these regions based on the match between the guide RNA (gRNA) and target DNA sequence. This study points to the off-target impact of mismatches between the gRNA and target DNA on exon regions of the TP53 gene, which are involved in regulating multiple genes and cellular functions. Off-target positions are typically evaluated using scoring methods. In this study, we have used latent class analysis to reveal subclasses of off-target positions. Thus, we have created the levels of off-target positions and evaluated the effects of mismatching positions within these classes using machine learning classifiers. The results revealed that mismatching positions could be categorized into three levels: low, middle, and high off-target positions. We have improved a computational framework to minimize off-target effects and to identify the PAM sequences in the gRNA design. Thus, carefully designed gRNAs will ensure that desired genetic edits are performed and target variants are achieved. This work will avail the future research aimed at optimizing genome editing by customizing CRISPR-Cas9 to target specific protospacer DNA through gRNA.

CRISPR-Cas Systems↗

Genetic mapping and predictive modeling of paralog synthetic lethality.

Paralogs are abundant in the human genome and thought to be a primary source of synthetic lethality, yet the vast paralogome remains largely uncharacterized. A digenic screen of 36,648 paralogous pairs in the human genome revealed that synthetic lethalities were infrequent and varied in penetrance in different tumor backgrounds. We hypothesized that the variable penetrance of synthetic lethalities resulted from complex polygenic interactions with different cellular contexts. A machine learning classifier of a subset of paralog pairs tested across 49 cancer models revealed that endogenous perturbations in related pathways predicted paralog synthetic lethality. Further, predictive modeling of paralog synthetic lethality showed that the strength of synthetic lethal interactions was largely due to the overlap and essentiality of the protein-protein interaction networks shared by the paralog pairs. Collectively, this study tested 36,648 digenic paralog interactions and delineated the key feature classes that underlie the heterogeneity of paralog synthetic lethalities.

Humans↗

Identification and analysis of deleterious human SNPs.

We have developed two methods of identifying which non-synonomous single base changes have a deleterious effect on protein function in vivo. One method, described elsewhere, analyzes the effect of the resulting amino acid change on protein stability, utilizing structural information. The other method, introduced here, makes use of the conservation and type of residues observed at a base change position within a protein family. A machine learning technique, the support vector machine, is trained on single amino acid changes that cause monogenic disease, with a control set of amino acid changes fixed between species. Both methods are used to identify deleterious single nucleotide polymorphisms (SNPs) in the human population. After carefully controlling for errors, we find that approximately one quarter of known non-synonymous SNPs are deleterious by these criteria, providing a set of possible contributors to human complex disease traits.

Animals↗

Identifying genes related to drug anticancer mechanisms using support vector machine.

In an effort to identify genes related to the cell line chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine was used to label genes into any of the five predefined anticancer drug mechanistic categories. Among dozens of unequivocally categorized genes, many were known to be causally related to the drug mechanisms. For example, a few genes were found to be involved in the biological process triggered by the drugs (e.g. DNA polymerase epsilon was the direct target for the drugs from DNA antimetabolites category). DNA repair-related genes were found to be enriched for about eight-fold in the resulting gene set relative to the entire gene set. Some uncharacterized transcripts might be of interest in future studies. This method of correlating the drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Antineoplastic Agents↗

Predicting essential genes in fungal genomes.

Essential genes are required for an organism's viability, and the ability to identify these genes in pathogens is crucial to directed drug development. Predicting essential genes through computational methods is appealing because it circumvents expensive and difficult experimental screens. Most such prediction is based on homology mapping to experimentally verified essential genes in model organisms. We present here a different approach, one that relies exclusively on sequence features of a gene to estimate essentiality and offers a promising way to identify essential genes in unstudied or uncultured organisms. We identified 14 characteristic sequence features potentially associated with essentiality, such as localization signals, codon adaptation, GC content, and overall hydrophobicity. Using the well-characterized baker's yeast Saccharomyces cerevisiae, we employed a simple Bayesian framework to measure the correlation of each of these features with essentiality. We then employed the 14 features to learn the parameters of a machine learning classifier capable of predicting essential genes. We trained our classifier on known essential genes in S. cerevisiae and applied it to the closely related and relatively unstudied yeast Saccharomyces mikatae. We assessed predictive success in two ways: First, we compared all of our predictions with those generated by homology mapping between these two species. Second, we verified a subset of our predictions with eight in vivo knockouts in S. mikatae, and we present here the first experimentally confirmed essential genes in this species.

Computational Biology↗

Comprehensive bioinformatic analysis of the specificity of human immunodeficiency virus type 1 protease.

Rapidly developing viral resistance to licensed human immunodeficiency virus type 1 (HIV-1) protease inhibitors is an increasing problem in the treatment of HIV-infected individuals and AIDS patients. A rational design of more effective protease inhibitors and discovery of potential biological substrates for the HIV-1 protease require accurate models for protease cleavage specificity. In this study, several popular bioinformatic machine learning methods, including support vector machines and artificial neural networks, were used to analyze the specificity of the HIV-1 protease. A new, extensive data set (746 peptides that have been experimentally tested for cleavage by the HIV-1 protease) was compiled, and the data were used to construct different classifiers that predicted whether the protease would cleave a given peptide substrate or not. The best predictor was a nonlinear predictor using two physicochemical parameters (hydrophobicity, or alternatively polarity, and size) for the amino acids, indicating that these properties are the key features recognized by the HIV-1 protease. The present in silico study provides new and important insights into the workings of the HIV-1 protease at the molecular level, supporting the recent hypothesis that the protease primarily recognizes a conformation rather than a specific amino acid sequence. Furthermore, we demonstrate that the presence of 1 to 2 lysine residues near the cleavage site of octameric peptide substrates seems to prevent cleavage efficiently, suggesting that this positively charged amino acid plays an important role in hindering the activity of the HIV-1 protease.

Algorithms↗

Symbiogenesis in learning classifier systems.

Symbiosis is the phenomenon in which organisms of different species live together in close association, resulting in a raised level of fitness for one or more of the organisms. Symbiogenesis is the name given to the process by which symbiotic partners combine and unify, that is, become genetically linked, giving rise to new morphologies and physiologies evolutionarily more advanced than their constituents. The importance of this process in the evolution of complexity is now well established. Learning classifier systems are a machine learning technique that uses both evolutionary computing techniques and reinforcement learning to develop a population of cooperative rules to solve a given task. In this article we examine the use of symbiogenesis within the classifier system rule base to improve their performance. Results show that incorporating simple rule linkage does not give any benefits. The concept of (temporal) encapsulation is then added to the symbiotic rules and shown to improve performance in ambiguous/non-Markov environments.

Algorithms↗

Identifying genes related to chemosensitivity using support vector machine.

In an effort to identify genes involved in chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine (SVM) is used to associate genes with any of five predefined anticancer drug mechanistic categories. The drug activity profiles are used as training examples to train the SVM and then the gene expression profiles are used as test examples to predict their associated mechanistic categories. This method of correlating drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Algorithms↗

Identification of Immune Response-Related Proteomic Biomarkers in Moyamoya Disease Using Serum Olink Proteomics.

Moyamoya disease, a rare chronic cerebrovascular disorder, requires invasive digital subtraction angiography (DSA) for diagnosis. This study employed high-throughput proteomics to identify plasma biomarkers for Moyamoya disease diagnosis. We conducted immunopanel analysis using the Olink platform to evaluate 92 immune-related proteins in plasma samples from 88 Moyamoya disease patients and 88 healthy controls. Key proteins were identified through differential expression analysis, GO, and KEGG enrichment analysis. A diagnostic model was constructed using LASSO regression, Boruta algorithm, and machine learning models including random forest and XGBoost. Validation of these proteins was performed using GEO external data sets, followed by prediction of potential therapeutic drugs and molecular docking validation through pharmacogenomic databases. A total of 44 differentially expressed proteins were identified through the Olink immunopanel, with 12 downregulated and 32 upregulated. GO and KEGG analyses revealed significant enrichment of these proteins in innate immune responses and signaling pathways such as NF-kB and MAPK. Through LASSO, random forest, and protein under-area analysis, four potential biomarkers for Moyamoya disease (MGMT, SIT1, PRDX1, TRAF2) were identified. A diagnostic model using these proteins showed the highest AUC value with the XGBoost model. Additionally, TRAF2 and PRDX1 exhibited significant expression differences in Moyamoya disease patients within the GEO data set. Our study revealed the immune landscape of Moyamoya disease, identified four biomarkers, and established a variety of diagnostic models.

Humans↗

Synthetic DNA barcodes identify singlets in scRNA-seq datasets and evaluate doublet algorithms.

Single-cell RNA sequencing (scRNA-seq) datasets contain true single cells, or singlets, in addition to cells that coalesce during the protocol, or doublets. Identifying singlets with high fidelity in scRNA-seq is necessary to avoid false negative and false positive discoveries. Although several methodologies have been proposed, they are typically tested on highly heterogeneous datasets and lack a priori knowledge of true singlets. Here, we leveraged datasets with synthetically introduced DNA barcodes for a hitherto unexplored application: to extract ground-truth singlets. We demonstrated the feasibility of our framework, "singletCode," to evaluate existing doublet detection methods across a range of contexts. We also leveraged our ground-truth singlets to train a proof-of-concept machine learning classifier, which outperformed other doublet detection algorithms. Our integrative framework can identify ground-truth singlets and enable robust doublet detection in non-barcoded datasets.

Algorithms↗

Support vector machine classification of 18F-FDG PET scans across subtypes of amyotrophic lateral sclerosis.

PURPOSE: While 18F-FDG PET imaging has demonstrated diagnostic value in people with Amyotrophic Lateral Sclerosis (PwALS) and group-level differences were identified between different disease subtypes (e.g., genetic and clinical variants), refining and validating a machine-learning-based subject-level diagnostic algorithm may improve the general applicability and reliability of 18F-FDG PET as a diagnostic tool in ALS. In this study, we employed support vector machines (SVM) to further explore the diagnostic potential of 18F-FDG PET in ALS, alongside its ability to classify between different genetic subtypes or clinical phenotypes. METHODS: 18F-FDG PET data of 36 healthy volunteers (HV), 25 people with ALS-mimicking diseases (Mimics), and 167 PwALS, grouped by genetic status (e.g., sporadic (sALS) or carrying a C9orf72 hexanucleotide repeat expansion (ALSC9orf72RE) and onset (bulbar or spinal) type, acquired with Biograph 'TruePoint' PET/CT scanner, were included in the study (Dataset 1). A second dataset of 183 PwALS and 31 Mimics acquired with Biograph 'HiRez' scanner was included as an independent cross-validation set (Dataset 2). PET images were spatially normalised to MNI space to fit linear SVMs with cross-validation. Only age-matched groups were considered to eliminate age-related effects. RESULTS: For Dataset 1, the linear SVM resulted in an average accuracy of 0.86 for the classification of ALS vs. HV, 0.53 for ALS vs. Mimics, 0.83 for ALSC9orf72RE vs. sALS, and 0.58 for bulbar vs. spinal onset. These findings were corroborated with Dataset2, with an accuracy of up to 0.76 for ALSC9orf72RE vs. sALS, and 0.59 for bulbar vs. spinal. CONCLUSION: 18F-FDG brain PET imaging, combined with SVM and age-matching, can distinguish between ALSC9orf72RE and sALS with good accuracy, but lacks sufficient discriminative power to differentiate between ALS and Mimics and between different sites of onset.

Humans↗