Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

An eco-informatics tool for microbial community studies: supervised classification of Amplicon Length Heterogeneity (ALH) profiles of 16S rRNA.

Support vector machines (SVM) and K-nearest neighbors (KNN) are two computational machine learning tools that perform supervised classification. This paper presents a novel application of such supervised analytical tools for microbial community profiling and to distinguish patterning among ecosystems. Amplicon length heterogeneity (ALH) profiles from several hypervariable regions of 16S rRNA gene of eubacterial communities from Idaho agricultural soil samples and from Chesapeake Bay marsh sediments were separately analyzed. The profiles from all available hypervariable regions were concatenated to obtain a combined profile, which was then provided to the SVM and KNN classifiers. Each profile was labeled with information about the location or time of its sampling. We hypothesized that after a learning phase using feature vectors from labeled ALH profiles, both these classifiers would have the capacity to predict the labels of previously unseen samples. The resulting classifiers were able to predict the labels of the Idaho soil samples with high accuracy. The classifiers were less accurate for the classification of the Chesapeake Bay sediments suggesting greater similarity within the Bay's microbial community patterns in the sampled sites. The profiles obtained from the V1+V2 region were more informative than that obtained from any other single region. However, combining them with profiles from the V1 region (with or without the profiles from the V3 region) resulted in the most accurate classification of the samples. The addition of profiles from the V 9 region appeared to confound the classifiers. Our results show that SVM and KNN classifiers can be effectively applied to distinguish between eubacterial community patterns from different ecosystems based only on their ALH profiles.

Artificial Intelligence↗

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models↗

Data-driven approaches in green microbiology: strategies for plant growth-promoting bacteria.

Plant growth-promoting bacteria (PGPB) are gaining attention as scalable biological solutions to enhance crop productivity and resilience. However, accurately identifying and characterizing PGPB remains challenging, particularly under variable environmental conditions where microbial functions are context-dependent and shaped by complex plant-microbe interactions. Advances in high-throughput sequencing have shifted the field from culture-dependent approaches to genome-informed strategies, enabling large-scale taxonomic and functional profiling. Although trait-based databases support the prediction of plant-beneficial genes, they capture only a fraction of the underlying biological complexity and often require labor-intensive analyses. Machine learning (ML) and deep learning (DL) have emerged as powerful tools to integrate genomic, physiological, and ecological data, enabling the prioritization of candidate strains with plant growth-promoting potential. To evaluate advances in the field, we conducted a systematic review of studies integrating ML and DL with PGPB characterization, assessing algorithm selection, performance, and target plant systems. Across 248 observations, only 6.0% of studies directly addressed PGPB screening, whereas the majority (77.4%) focused on plant disease detection, revealing a substantial gap in the application of AI to beneficial microorganisms for plant growth. Convolutional neural networks (CNNs) were the most frequently applied algorithms, largely driven by image-based phenotyping tasks. Overall, the field is constrained by limited datasets, high computational demands, and challenges in modeling multispecies and host-associated interactions. We highlight the need for integrative and interpretable ML and DL frameworks that bridge genomic data and functional validation. Such approaches represent a promising path toward scalable, data-driven discovery and deployment of bioinoculants in sustainable agriculture.

Agriculture↗

A wiring of the human nucleolus.

Recent proteomic efforts have created an extensive inventory of the human nucleolar proteome. However, approximately 30% of the identified proteins lack functional annotation. We present an approach of assigning function to uncharacterized nucleolar proteins by data integration coupled to a machine-learning method. By assembling protein complexes, we present a first draft of the human ribosome biogenesis pathway encompassing 74 proteins and hereby assign function to 49 previously uncharacterized proteins. Moreover, the functional diversity of the nucleolus is underlined by the identification of a number of protein complexes with functions beyond ribosome biogenesis. Finally, we were able to obtain experimental evidence of nucleolar localization of 11 proteins, which were predicted by our platform to be associates of nucleolar complexes. We believe other biological organelles or systems could be "wired" in a similar fashion, integrating different types of data with high-throughput proteomics, followed by a detailed biological analysis and experimental validation.

Artificial Intelligence↗

dsRNAscan maps human dsRNAome, revealing conservation, intermolecular dsRNA, and correlates of ADAR dependency.

The human transcriptome contains millions of A-to-I editing sites arising from an unclear number of poorly characterized dsRNAs. Editing sites reveal the presence of dsRNA, but this method is limited by transcription levels, read depth, and ADAR expression and cannot identify unedited dsRNA. To address these limitations, we developed dsRNAscan. Applying dsRNAscan to the human genome predicted 5 million dsRNAs, mostly in repetitive and intergenic regions. Machine learning models trained on A-to-I editing and RNA structure-probing data identified ∼2.4 million high-confidence predictions, which were enriched at dsRNA-binding protein binding sites. Additionally, we predicted hundreds of dsRNAs conserved across vertebrates and observed thousands of editing-enriched regions suspected to arise from intermolecular dsRNAs formed with sense-antisense transcripts. Quantifying expression of intramolecular and intermolecular dsRNAs accessible to cytoplasmic immune sensors revealed that their ratio correlated with ADAR dependency across cancer cell lines. The human dsRNAome is available as a resource at https://dsrna.chpc.utah.edu/.

A-to-I RNA editing↗

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO↗

Effectiveness of artificial intelligence in nursing simulation education: A systematic review, meta-analysis and bibliometric visualization analysis.

OBJECTIVES: To synthesize the roles and core functions of AI in nursing simulation education for nursing students via systematic review, quantitatively evaluate its effects on students' knowledge and skill outcomes through meta-analysis, and map the research landscape and development trends of this field through bibliometric visualization analysis. DESIGN: Systematic review, meta-analysis and bibliometric visualization analysis. DATA SOURCES: Eight electronic databases: PubMed, Web of Science, MEDLINE, ERIC, Academic Search Complete, China National Knowledge Infrastructure (CNKI), Wanfang Database, VIP Chinese Science and Technology Journal Database (VIP) were employed to search studies from the time of construction to 16 December 2025. REVIEW METHODS: Studies meeting the inclusion criteria were screened. The revised Cochrane Risk of Bias tool (ROB 2) and Joanna Briggs Institute (JBI) critical appraisal checklists were used for quality assessment. Meta-analysis was performed with Review Manager 5.4, and bibliometric visualization analysis was conducted using VOSviewer 1.6.20 and Bibliometrix (based on R4.4.3). RESULTS: A total of 61 studies were included. AI primarily played two roles in nursing simulation education: peer-type new subject (n = 24) and direct mediator (n = 22). Meta-analysis showed that AI interventions significantly improved nursing students' knowledge (SMD = 1.49, 95% CI [0.55,2.43], p = 0.002) and skills (SMD = 0.66, 95% CI [0.02,1.31], p = 0.04). Bibliometric analysis identified that the United States of America and China were the two main contributing countries in this field, and the key motor themes included generative artificial intelligence, virtual patients, and geriatric care. CONCLUSIONS: AI exerts positive effects on nursing students' knowledge acquisition and skill enhancement in simulation education, with peer-type new subject and direct mediator as the dominant roles. Future research should focus on expanding AI applications in multi-specialty simulation scenarios, activating the data-driven value of machine learning, and strengthening international collaboration and standardization construction, so as to promote the sustainable development of AI-integrated nursing simulation education.

Humans↗

On the relationship between deterministic and probabilistic directed Graphical models: from Bayesian networks to recursive neural networks.

Machine learning methods that can handle variable-size structured data such as sequences and graphs include Bayesian networks (BNs) and Recursive Neural Networks (RNNs). In both classes of models, the data is modeled using a set of observed and hidden variables associated with the nodes of a directed acyclic graph. In BNs, the conditional relationships between parent and child variables are probabilistic, whereas in RNNs they are deterministic and parameterized by neural networks. Here, we study the formal relationship between both classes of models and show that when the source nodes variables are observed, RNNs can be viewed as limits, both in distribution and probability, of BNs with local conditional distributions that have vanishing covariance matrices and converge to delta functions. Conditions for uniform convergence are also given together with an analysis of the behavior and exactness of Belief Propagation (BP) in 'deterministic' BNs. Implications for the design of mixed architectures and the corresponding inference algorithms are briefly discussed.

Bayes Theorem↗

Pipelining of Fuzzy ARTMAP without matchtracking: correctness, performance bound, and Beowulf evaluation.

Fuzzy ARTMAP neural networks have been proven to be good classifiers on a variety of classification problems. However, the time that Fuzzy ARTMAP takes to converge to a solution increases rapidly as the number of patterns used for training is increased. In this paper we examine the time Fuzzy ARTMAP takes to converge to a solution and we propose a coarse grain parallelization technique, based on a pipeline approach, to speed-up the training process. In particular, we have parallelized Fuzzy ARTMAP without the match-tracking mechanism. We provide a series of theorems and associated proofs that show the characteristics of Fuzzy ARTMAP's, without matchtracking, parallel implementation. Results run on a BEOWULF cluster with three large databases show linear speedup as a function of the number of processors used in the pipeline. The databases used for our experiments are the Forrest CoverType database from the UCI Machine Learning repository and two artificial databases, where the data generated were 16-dimensional Gaussian distributed data belonging to two distinct classes, with different amounts of overlap (5% and 15%).

Databases as Topic↗

Classification of fMRI independent components using IC-fingerprints and support vector machine classifiers.

We present a general method for the classification of independent components (ICs) extracted from functional MRI (fMRI) data sets. The method consists of two steps. In the first step, each fMRI-IC is associated with an IC-fingerprint, i.e., a representation of the component in a multidimensional space of parameters. These parameters are post hoc estimates of global properties of the ICs and are largely independent of a specific experimental design and stimulus timing. In the second step a machine learning algorithm automatically separates the IC-fingerprints into six general classes after preliminary training performed on a small subset of expert-labeled components. We illustrate this approach in a multisubject fMRI study employing visual structure-from-motion stimuli encoding faces and control random shapes. We show that: (1) IC-fingerprints are a valuable tool for the inspection, characterization and selection of fMRI-ICs and (2) automatic classifications of fMRI-ICs in new subjects present a high correspondence with those obtained by expert visual inspection of the components. Importantly, our classification procedure highlights several neurophysiologically interesting processes. The most intriguing of which is reflected, with high intra- and inter-subject reproducibility, in one IC exhibiting a transiently task-related activation in the 'face' region of the primary sensorimotor cortex. This suggests that in addition to or as part of the mirror system, somatotopic regions of the sensorimotor cortex are involved in disambiguating the perception of a moving body part. Finally, we show that the same classification algorithm can be successfully applied, without re-training, to fMRI collected using acquisition parameters, stimulation modality and timing considerably different from those used for training.

Algorithms↗

Bioinformatics in proteomics: application, terminology, and pitfalls.

Bioinformatics applies data mining, i.e., modern computer-based statistics, to biomedical data. It leverages on machine learning approaches, such as artificial neural networks, decision trees and clustering algorithms, and is ideally suited for handling huge data amounts. In this article, we review the analysis of mass spectrometry data in proteomics, starting with common pre-processing steps and using single decision trees and decision tree ensembles for classification. Special emphasis is put on the pitfall of overfitting, i.e., of generating too complex single decision trees. Finally, we discuss the pros and cons of the two different decision tree usages.

Computational Biology↗

MULTIPREVENT: Integrated screening for smoking-related multimorbidity using low-dose chest computed tomography.

OBJECTIVES: Tobacco consumption, combined with individual genetic predispositions, contributes to an age-dependent risk not only for lung cancer but also for other non-communicable diseases (NCDs) such as cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), osteoporosis, and diabetes. The MULTIPREVENT project aims to validate whether low-dose computed tomography (LDCT) of the chest, combined with simple biomarkers, functional tests, and genomic profiling, can serve as an effective tool for comprehensive health assessment and risk prediction of multimorbidity in adults. STUDY DESIGN: The study is based on a prospective epidemiological design involving 3000 participants from the MOLTEST-BIS lung cancer screening cohort (2016-2018). These participants, aged 50-79 years (during MOLTEST-BIS) and with a smoking history of at least 30 pack-years, will undergo two follow-up assessments in 2025-2027 and 2030-2032. METHODS: Each follow-up includes LDCT, spirometry, standardized blood pressure measurement, anthropometric evaluation, biomarker assessment (lipid profile, lipoprotein(a), glycated haemoglobin), and health-related questionnaires. Genetic profiling will be performed using the Illumina Infinium Global Screening Arrays approach to identify inherited predispositions to major NCDs. All data, clinical, imaging (including radiomics), molecular, and genetic, will be integrated through machine learning algorithms to develop AI-based risk prediction models. RESULTS: The MULTIPREVENT study is expected to generate a wide range of scientific, clinical, and infrastructural results that will serve as a foundation for future public health initiatives in integrated prevention. CONCLUSIONS: By linking imaging and biochemical markers, genetic susceptibility, and clinical parameters within a longitudinal design, MULTIPREVENT will establish data-driven, AI-supported prevention strategies aimed at reducing morbidity and mortality among adults exposed to tobacco. The project will also serve as a model for population-based multimorbidity prevention programs.

Humans↗

Crosstalk between cysteine and lysine modifications: Integrating redox and metabolic regulation.

Protein post-translational modifications (PTMs) on amino acid residues enable dynamic cellular responses to changes in metabolic and redox state. Cysteine and lysine are among the most extensively modified amino acid residues, with both undergoing a diversity of acylation and oxidative modifications. Indeed, proximal (<10&#x202f;&#xc5;) cysteine and lysine residues may form integration nodes for crosstalk between metabolism and redox homeostasis pathways. This review highlights the interaction of proximal Cys-Lys residues, including influence on residue pKa by local electrostatics, cysteine-to-lysine transfer of PTM moieties, and covalent crosslinking. We discuss candidate Cys-Lys regulatory pairs in proteins involved in redox regulation, proteostasis, metabolic adaptation and inflammation. We further utilize computational modeling to identify proximity between cysteine and lysine residues in proteins known to be regulated by acylation and oxidative PTMs, and to demonstrate changes in these distances and local electrostatic potential due to lysine acetylation. Finally, we review how mass spectrometry-based proteomics and machine-learning PTM predictive tools can enable the identification, validation, and interpretation of proximal Cys-Lys interactions that regulate cellular responses to oxidative challenge and metabolic flux.

Cysteine↗

Precision medicine in combating antimicrobial resistance: A comprehensive review.

Antimicrobial resistance (AMR) represents one of the most pressing threats to global public health, undermining the effectiveness of modern antimicrobial therapy and challenging decades of medical progress. This comprehensive review examines the transition from broad-spectrum empirical therapy toward precision medicine as an integrated framework for improving antimicrobial use and combating AMR. Precision medicine seeks to tailor treatment decisions by combining pathogen-specific genomic and resistance data with relevant host characteristics to optimize therapy while limiting unnecessary antimicrobial exposure and the selective pressures that drive resistance. The review synthesizes advances reported from 2020, highlighting established and emerging approaches including rapid molecular diagnostics, next-generation sequencing, CRISPR-based detection, machine learning (ML)-assisted decision support, precision dosing, and targeted therapeutics such as bacteriophage therapy, antimicrobial peptides, and bacterial proteolysis-targeting chimeras. Rather than functioning as isolated technologies, these approaches achieve their greatest clinical value when integrated within antimicrobial stewardship programs and a One Health framework that recognizes the interconnected human, animal, and environmental drivers of resistance. Despite considerable progress, important challenges remain, including equitable access to advanced technologies, interpretation of increasingly complex datasets, workforce and infrastructure limitations, and evolving regulatory pathways for novel diagnostics and therapeutics. This review concludes that while precision medicine is not a standalone solution, its successful implementation will depend on coordinated integration of diagnostics, host factors, computational tools, pharmacological optimization, and stewardship strategies to improve patient outcomes while preserving the long-term effectiveness of existing antimicrobials.

Antimicrobial resistance↗

Polygenic risk scores for rheumatoid arthritis and idiopathic pulmonary fibrosis and associations with RA, interstitial lung abnormalities, and quantitative interstitial abnormalities among smokers.

OBJECTIVE: Genome-wide association studies (GWAS) facilitate construction of polygenic risk scores (PRSs) for rheumatoid arthritis (RA) and idiopathic pulmonary fibrosis (IPF). We investigated associations of RA and IPF PRSs with RA and high-resolution chest computed tomography (HRCT) parenchymal lung abnormalities. METHODS: Participants in COPDGene, a prospective multicenter cohort of current/former smokers, had chest HRCT at study enrollment. Using genome-wide genotyping, RA and IPF PRSs were constructed using GWAS summary statistics. HRCT imaging underwent visual inspection for interstitial lung abnormalities (ILA) and quantitative CT (QCT) analysis using a machine-learning algorithm that quantified percentage of normal lung, interstitial abnormalities, and emphysema. RA was identified through self-report and DMARD use. We investigated associations of RA and IPF PRSs with RA, ILA, and QCT features using multivariable logistic and linear regression. RESULTS: We analyzed 9,230 COPDGene participants (mean age 59.6 years, 46.4 % female, 67.2 % non-Hispanic White, 32.8 % Black/African American). In non-Hispanic White participants, RA PRS was associated with RA diagnosis (OR 1.32 per unit, 95 %CI 1.18-1.49) but not ILA or QCT features. Among non-Hispanic White participants, IPF PRS was associated with ILA (OR 1.88 per unit, 95 %CI 1.52-2.32) and quantitative interstitial abnormalities (adjusted &#x3b2;=+0.50 % per unit, p = 7.3 &#xd7; 10-8) but not RA. There were no statistically significant associations among Black/African American participants. CONCLUSIONS: RA and IPF PRSs were associated with their intended phenotypes among non-Hispanic White participants but performed poorly among Black/African American participants. PRS may have future application to risk stratify for RA diagnosis among patients with ILD or for ILD among patients with RA.

Humans↗

Unveiling tumor heterogeneity by single cell RNA-sequencing: From basic considerations to clinical applications.

Tumor heterogeneity-encompassing diverse cellular phenotypes, genomic alterations, and microenvironmental contexts-is a principal barrier to effective cancer therapy. Single-cell RNA sequencing (scRNA-seq) has transformed our ability to resolve this complexity by capturing transcriptomes at single-cell resolution. Here, we review the technical foundations required for high-quality scRNA-seq studies. We then trace the evolution of scRNA-seq platforms from manual micromanipulation to high-throughput systems, and describe the computational pipelines that enable reliable data interpretation. The application of scRNA-seq is exemplarily shown in the context of lung cancer, where single-cell profiling has revealed (i) the clonal and sub-clonal architecture of tumors, (ii) extensive remodeling of the immune microenvironment, iii) key mechanisms underlying resistance to targeted agents and immune-checkpoint blockade, and (iv) the dynamics of neo-antigen-specific T-cell responses. Integrating machine-learning techniques-such as deep-learning classifiers and graph-based models-with single-cell transcriptomic data has markedly sped up biomarker discovery, produced more accurate risk-stratification scores, and enabled the generation of patient-specific therapeutic predictions. We surveyed the major trial registry ClinicalTrials.gov and identified &#x223c;380&#xa0;ongoing or completed studies that explicitly incorporate scRNA-seq as a correlative or pharmacodynamic endpoint. Overall, the analysis shows that scRNA-seq becomes an increasingly important component of modern trials, providing high-resolution cellular and molecular readouts that complement conventional imaging and bulk-omics endpoints. While key challenges remain, ranging from costs, scalability and need for rigorous validation before routine clinical deployment, ongoing technological advances continue to expand the potential of scRNA-seq as a cornerstone of precision medicine.

Humans↗

Molecular signature of primate astrocytes reveals pathways and regulatory changes contributing to human brain evolution.

Astrocytes contribute to the development and regulation of the higher-level functions of the brain, the critical targets of evolution. However, how astrocytes evolve in primates is unsettled. Here, we obtain human, chimpanzee, and macaque induced pluripotent stem-cell-derived astrocytes (iAstrocytes). Human iAstrocytes are bigger and more complex than the non-human primate iAstrocytes. We identify new loci contributing to the increased human astrocyte. We show that genes and pathways implicated in long-range intercellular signaling are activated in the human iAstrocytes and partake in controlling iAstrocyte complexity. Genes downregulated in human iAstrocytes frequently relate to neurological disorders and were decreased in adult brain samples. Through regulome analysis and machine learning, we uncover that functional activation of enhancers coincides with a previously unappreciated, pervasive gain of "stripe" transcription factor binding sites. Altogether, we reveal the transcriptomic signature of primate astrocyte evolution and a mechanism driving the acquisition of the regulatory potential of enhancers.

Astrocytes↗

Intelligent software for laboratory automation.

The automation of laboratory techniques has greatly increased the number of experiments that can be carried out in the chemical and biological sciences. Until recently, this automation has focused primarily on improving hardware. Here we argue that future advances will concentrate on intelligent software to integrate physical experimentation and results analysis with hypothesis formulation and experiment planning. To illustrate our thesis, we describe the 'Robot Scientist' - the first physically implemented example of such a closed loop system. In the Robot Scientist, experimentation is performed by a laboratory robot, hypotheses concerning the results are generated by machine learning and experiments are allocated and selected by a combination of techniques derived from artificial intelligence research. The performance of the Robot Scientist has been evaluated by a rediscovery task based on yeast functional genomics. The Robot Scientist is proof that the integration of programmable laboratory hardware and intelligent software can be used to develop increasingly automated laboratories.

Algorithms↗