Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

A case study where biology inspired a solution to a computer science problem.

This paper describes how the biological theory of gene duplication described in Susumu Ohno's provocative book, Evolution by Means of Gene Duplication, was brought to bear on a vexatious problem from the domain of automated machine learning, namely the problem of architecture discovery. Six new architecture-altering operations for genetic programming were motivated by the way that new biological structures, functions, and behaviors arise in nature using gene duplication. Genetic programming with the new architecture-altering operations was then applied to the transmembrane protein segment identification problem. The out-of-sample error rate for the best genetically-evolved program achieved was slightly better than that of previously-reported human-written algorithms for this problem.

Amino Acid Sequence↗

Identification of divergent functions in homologous proteins by induction over conserved modules.

Homologous proteins do not necessarily exhibit identical biochemical function. Despite this fact, local or global sequence similarity is widely used as an indication of functional identity. Of the 1327 Enzyme Commission defined functional classes with more than one annotated example in the sequence databases, similarity scores alone are inadequate in 251 (19%) of the cases. We test the hypothesis that conserved domains, as defined in the ProDom database, can be used to discriminate between alternative functions for homologous proteins in these cases. Using machine learning methods, we were able to induce correct discriminators for more than half of these 251 challenging functional classes. These results show that the combination of modular representations of proteins with sequence similarity improves the ability to infer function from sequence over similarity scores alone.

Alcohol Dehydrogenase↗

Active learning with support vector machines in the drug discovery process.

We investigate the following data mining problem from computer-aided drug design: From a large collection of compounds, find those that bind to a target molecule in as few iterations of biochemical testing as possible. In each iteration a comparatively small batch of compounds is screened for binding activity toward this target. We employed the so-called "active learning paradigm" from Machine Learning for selecting the successive batches. Our main selection strategy is based on the maximum margin hyperplane-generated by "Support Vector Machines". This hyperplane separates the current set of active from the inactive compounds and has the largest possible distance from any labeled compound. We perform a thorough comparative study of various other selection strategies on data sets provided by DuPont Pharmaceuticals and show that the strategies based on the maximum margin hyperplane clearly outperform the simpler ones.

Computer-Aided Design↗

Learning in brains and machines.

The problem of learning is arguably at the very core of the problem of intelligence, both biological and artificial. In this paper we sketch some of our work over the last ten years in the area of supervised learning, focusing on three interlinked directions of research: theory, engineering applications (that is, making intelligent software) and neuroscience (that is, understanding the brain's mechanisms of learning).

Brain↗

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence↗

Spatiotemporal genomic analysis and risk assessment of the plasmids carrying blaOXA-48-like genes based on a large-scale international dataset.

BACKGROUND: The spread of OXA-48-like carbapenemases represents a major public health challenge. Although previous studies have investigated OXA-48-like carbapenemases risk factors, nosocomial dissemination, and plasmid dynamics, an integrated plasmid-centered framework combining complete plasmid mining, transmission-unit analysis, phylogenetic reconstruction, and machine learning-based risk assessment remains limited. METHODS: We systematically collected 747 complete plasmid sequences carrying blaOXA-48-like genes from the NCBI database, establishing the largest collections of complete plasmid sequences to date. Using an integrative framework of population genomics, phylogenetic dating, and machine learning, this study aimed to characterize the dissemination patterns, plasmid replicon diversity, transmission units, mobile genetic elements, co-resistance profiles, and risk classification of these plasmid. RESULTS: Plasmids carrying blaOXA-48-like genes were detected across 50 countries on six continents, with blaOXA-48 predominating in Europe, blaOXA-181 in South Asia, and blaOXA-232 largely in Asia. IncL and ColKP3/IncX3 replicons, together with Tn1999.2 and other MGEs, were central drivers of plasmid maintenance and spread. Sixteen transmission units were defined, with AA068_Cluster3 estimated to have originated in the Netherlands around 2005 before expanding to Europe, the Middle East, Asia, and North America. Co-resistance analyses revealed frequent modules involving aminoglycoside and quinolone resistance, with qnrS1 and aph(3'')-Ib most prevalent. Notably, high-risk transposon structures were often identified in non-clinical environments, underscoring their cross-ecological transmission potential. Machine learning-based classification models showed good internal performance for predefined composite-risk categories, with plasmid mobility, clinical/non-clinical source composition, and host background contributing to the classification results. CONCLUSIONS: This study provides a large-scale plasmid-centered genomic analysis of publicly available complete plasmid sequences carrying blaOXA-48-like genes, integrating transmission-unit inference, phylogeographic reconstruction, mobile genetic element and co-resistance profiling, and composite genomic risk stratification. This gene-centered framework may support future One Health-oriented antimicrobial resistance surveillance and prioritization of plasmids with higher dissemination and resistance potential.

Plasmids↗

Mutual learning in a tree parity machine and its application to cryptography.

Mutual learning of a pair of tree parity machines with continuous and discrete weight vectors is studied analytically. The analysis is based on a mapping procedure that maps the mutual learning in tree parity machines onto mutual learning in noisy perceptrons. The stationary solution of the mutual learning in the case of continuous tree parity machines depends on the learning rate where a phase transition from partial to full synchronization is observed. In the discrete case the learning process is based on a finite increment and a full synchronized state is achieved in a finite number of steps. The synchronization of discrete parity machines is introduced in order to construct an ephemeral key-exchange protocol. The dynamic learning of a third tree parity machine (an attacker) that tries to imitate one of the two machines while the two still update their weight vectors is also analyzed. In particular, the synchronization times of the naive attacker and the flipping attacker recently introduced in Ref. 9 are analyzed. All analytical results are found to be in good agreement with simulation results.

Journal Article↗

Generalization properties of finite-size polynomial support vector machines

The learning properties of finite-size polynomial support vector machines are analyzed in the case of realizable classification tasks. The normalization of the high-order features acts as a squeezing factor, introducing a strong anisotropy in the patterns distribution in feature space. As a function of the training set size, the corresponding generalization error presents a crossover, more or less abrupt depending on the distribution's anisotropy and on the task to be learned, between a fast-decreasing and a slowly decreasing regime. This behavior corresponds to the stepwise decrease found by Dietrich et al. [Phys. Rev. Lett. 82, 2975 (1999)] in the thermodynamic limit. The theoretical results are in excellent agreement with the numerical simulations.

Journal Article↗

Support vector machine classification on the web.

The support vector machine (SVM) learning algorithm has been widely applied in bioinformatics. We have developed a simple web interface to our implementation of the SVM algorithm, called Gist. This interface allows novice or occasional users to apply a sophisticated machine learning algorithm easily to their data. More advanced users can download the software and source code for local installation. The availability of these tools will permit more widespread application of this powerful learning algorithm in bioinformatics.

Algorithms↗

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi‑omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans↗

Relevance vector machine and support vector machine classifier analysis of scanning laser polarimetry retinal nerve fiber layer measurements.

PURPOSE: To classify healthy and glaucomatous eyes using relevance vector machine (RVM) and support vector machine (SVM) learning classifiers trained on retinal nerve fiber layer (RNFL) thickness measurements obtained by scanning laser polarimetry (SLP). METHODS: Seventy-two eyes of 72 healthy control subjects (average age = 64.3 +/- 8.8 years, visual field mean deviation = -0.71 +/- 1.2 dB) and 92 eyes of 92 patients with glaucoma (average age = 66.9 +/- 8.9 years, visual field mean deviation = -5.32 +/- 4.0 dB) were imaged with SLP with variable corneal compensation (GDx VCC; Laser Diagnostic Technologies, San Diego, CA). RVM and SVM learning classifiers were trained and tested on SLP-determined RNFL thickness measurements from 14 standard parameters and 64 sectors (approximately 5.6 degrees each) obtained in the circumpapillary area under the instrument-defined measurement ellipse (total 78 parameters). Ten-fold cross-validation was used to train and test RVM and SVM classifiers on unique subsets of the full 164-eye data set and areas under the receiver operating characteristic (AUROC) curve for the classification of eyes in the test set were generated. AUROC curve results from RVM and SVM were compared to those for 14 SLP software-generated global and regional RNFL thickness parameters. Also reported was the AUROC curve for the GDx VCC software-generated nerve fiber indicator (NFI). RESULTS: The AUROC curves for RVM and SVM were 0.90 and 0.91, respectively, and increased to 0.93 and 0.94 when the training sets were optimized with sequential forward and backward selection (resulting in reduced dimensional data sets). AUROC curves for optimized RVM and SVM were significantly larger than those for all individual SLP parameters. The AUROC curve for the NFI was 0.87. CONCLUSIONS: Results from RVM and SVM trained on SLP RNFL thickness measurements are similar and provide accurate classification of glaucomatous and healthy eyes. RVM may be preferable to SVM, because it provides a Bayesian-derived probability of glaucoma as an output. These results suggest that these machine learning classifiers show good potential for glaucoma diagnosis.

Aged↗

Prediction of cytochrome P450 3A4, 2D6, and 2C9 inhibitors and substrates by using support vector machines.

Statistical learning methods have been used in developing filters for predicting inhibitors of two P450 isoenzymes, CYP3A4 and CYP2D6. This work explores the use of different statistical learning methods for predicting inhibitors of these enzymes and an additional P450 enzyme, CYP2C9, and the substrates of the three P450 isoenzymes. Two consensus support vector machine (CSVM) methods, "positive majority" (PM-CSVM) and "positive probability" (PP-CSVM), were used in this work. These methods were first tested for the prediction of inhibitors of CYP3A4 and CYP2D6 by using a significantly higher number of inhibitors and noninhibitors than that used in earlier studies. They were then applied to the prediction of inhibitors of CYP2C9 and substrates of the three enzymes. Both methods predict inhibitors of CYP3A4 and CYP2D6 at a similar level of accuracy as those of earlier studies. For classification of inhibitors of CYP2C9, the best CSVM method gives an accuracy of 88.9% for inhibitors and 96.3% for noninhibitors. The accuracies for classification of substrates and nonsubstrates of CYP3A4, CYP2D6, and CYP2C9 are 98.2 and 90.9%, 96.6 and 94.4%, and 85.7 and 98.8%, respectively. Both CSVM methods are potentially useful as filters for predicting inhibitors and substrates of P450 isoenzymes. These methods generally give better accuracies than single SVM classification systems, and the performance of the PP-CSVM method is slightly better than that of the PM-CSVM method.

Algorithms↗

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis. A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST + AI for prediction model studies. Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST + AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection. AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans↗

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans↗

Representation learning for multi-modal spatially resolved transcriptomics data.

MOTIVATION: Spatial transcriptomics enables in-depth molecular characterization of samples on a morphology and RNA level while preserving spatial location. Integrating the resulting multi-modal data is an unsolved problem, and developing new solutions in precision medicine depends on improved methodologies. RESULTS: We introduce AESTETIK, a convolutional deep learning model that jointly integrates spatial, transcriptomics, and morphology information to learn accurate spot representations. AESTETIK yielded substantially improved cluster assignments on widely adopted technology platforms (e.g. 10x Genomics™, NanoString™) across multiple datasets. We achieved performance enhancement on structured tissues (e.g. brain) with a 21% increase in median ARI over previous state-of-the-art methods. Notably, AESTETIK also demonstrated superior performance on cancer tissues with heterogeneous cell populations, showing a 2-fold increase in breast cancer, 79% in melanoma, and 21% in liver cancer. We expect that these advances will enable a multi-modal understanding of key biological processes. AVAILABILITY AND IMPLEMENTATION: AESTETIK is implemented in Python 3 and is available as open source software at http://www.github.com/ratschlab/aestetik. The Snakemake pipeline for reproducing the results is available at http://www.github.com/ratschlab/st-rep.

Spatial Transcriptomics↗

Analysis of alcoholism data using support vector machines.

A supervised learning method, support vector machine, was used to analyze the microsatellite marker dataset of the Collaborative Study on the Genetics of Alcoholism Problem 1 for the Genetic Analysis Workshop 14. Twelve binary-valued phenotype variables were chosen for analyses using the markers from all autosomal chromosomes. Using various polynomial kernel functions of the support vector machine and randomly divided genome regions, we were able to observe the association of some marker sets with the chosen phenotypes and thus reduce the size of the dataset. The successful classifications established with the chosen support vector machine kernel function had high levels of correctness for each prediction, e.g., 96% in the fourfold cross-validations. However, owing to the limited sample data, we were not able to test the predictions of the classifiers in the new sample data.

Alcoholism↗