Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Predicting the efficiency of UAG translational stop signal through studies of physicochemical properties of its composite mono- and dinucleotides.

In this study, we explored the problem of predicting the UAG stop-codon read-through efficiency. The reported nucleotide sequences were first converted into physicochemical property vectors before being presented to a machine learning algorithm. Two sets of physicochemical properties were applied: one for mononucleosides (in terms of steric bulk, hydrophobicity and electronics) and another for dinucleotides. To the best of our knowledge, this is the first report of how dinucleotides are converted into principle components derived from NMR chemical shift data. A few efficiency prediction models were then derived and a comparison between mononucleoside and dinucleotide-based models was shown. In the derived models, the coefficients of these property based predictors lend themselves to bio-physical interpretations, an advantage which is demonstrated in this study via a prediction model based on the steric bulk factor. Although it is quite simple, the steric bulk factor model explained well the effect of sequence variations surrounding the amber stop codon and the tRNA bearing UCCU anticodon. We further proposed new alternatives at position -1 and +4 of a UAG stop codon sequence to enhance the readthrough efficiency. This research may contribute to a better understanding of the readthrough mechanisms and may also help to study the normal translation termination process.

Codon, Terminator↗

Retrieving definitional content for ontology development.

Ontology construction requires an understanding of the meaning and usage of its encoded concepts. While definitions found in dictionaries or glossaries may be adequate for many concepts, the actual usage in expert writing could be a better source of information for many others. The goal of this paper is to describe an automated procedure for finding definitional content in expert writing. The approach uses machine learning on phrasal features to learn when sentences in a book contain definitional content, as determined by their similarity to glossary definitions provided in the same book. The end result is not a concise definition of a given concept, but for each sentence, a predicted probability that it contains information relevant to a definition. The approach is evaluated automatically for terms with explicit definitions, and manually for terms with no available definition.

Bayes Theorem↗

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products↗

Integrative proteomics and bioinformatics pipelines for PTM profiling.

Post-translational modifications (PTMs) regulate protein function across all life forms and allow plants to respond rapidly to biotic and abiotic stress. Over 450 PTM types have been described across organisms, of which 23-33 have been experimentally confirmed in plants, including phosphorylation, acetylation, methylation, glycosylation, ubiquitination, and sumoylation. These modifications are highly dynamic and often reversible, and frequently act in combination, or "crosstalk," to fine-tune cellular processes. Advances in high-resolution mass spectrometry and large-scale genome sequencing continue to expand the catalogue of known PTM sites, while machine learning and deep learning approaches increasingly support prediction of PTM site localization and function. Unlike broader surveys of plant PTMs, this review focuses specifically on O-phosphorylation and Lys-N(ε)-acetylation, the two best-characterized and most extensively crosstalking PTMs in plants, and integrates four perspectives: the historical development of proteomic and bioinformatics approaches to these modifications; current mass spectrometry-based workflows and enrichment strategies; the bioinformatics tools and databases available for their analysis; and the technical and species-related challenges, particularly in non-model plants, that currently limit their study. We close by outlining priority directions for future research, including multi-omics integration, AI-based prediction, and the translation of PTM knowledge into crop stress resilience and breeding applications.

Protein Processing, Post-Translational↗

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape↗

Protein sequence-based risk classification for human papillomaviruses.

Human papillomaviruses (HPVs) are small DNA tumor viruses which infect epithelial tissues and induce hyperproliferative lesions. Infection by high-risk genital HPVs is associated with the development of anogenital cancers. Classification of risk types is important in understanding the mechanisms in infection and in developing novel instruments for medical examination such as DNA microarrays. The sequence-based classification methods are useful in classifying risk types by considering residues in conserved positions. In this paper, we present a machine learning approach to the classification of HPV risk types by using the protein sequences. Our approach is based on the hidden Markov model and the kernel method. The former searches informative subsequence positions and the latter computes efficiently to classify protein sequences. In the experiments, the classifier predicted four unknown HPV types exactly. An additional result shows that the kernel-based classifiers learned with more informative subsequences outperform the classifiers learned with the whole sequence or random subsequences.

Amino Acid Sequence↗

Building an ontology of adverse drug reactions for automated signal generation in pharmacovigilance.

Automated signal generation in pharmacovigilance implements unsupervised statistical machine learning techniques in order to discover unknown adverse drug reactions (ADR) in spontaneous reporting systems. The impact of the terminology used for coding ADRs has not been addressed previously. The Medical Dictionary for Regulatory Activities (MedDRA) used worldwide in pharmacovigilance cases does not provide formal definitions of terms. We have built an ontology of ADRs to describe semantics of MedDRA terms. Ontological subsumption and approximate matching inferences allow a better grouping of medically related conditions. Signal generation performances are significantly improved but time consumption related to modelization remains very important.

Adverse Drug Reaction Reporting Systems↗

Large-scale discovery platform enables identification of peptides targeting drug-resistant candidiasis.

Natural products have an unparalleled track record as sources of clinical drugs. Among them, nonribosomal peptides (NRPs) stand as one of the most therapeutically significant classes, encompassing numerous approved anti-infective and anticancer agents. Yet, discovering bioactive NRPs remains profoundly challenging due to their complex biosynthesis and chemical architecture. Here, we present NPDiscover, a pathogen-oriented, scalable bioinformatics platform that integrates genome mining, metabolomics, and machine learning to identify NRPs active against drug-resistant pathogens. Applying NPDiscover to Actinobacteria datasets, we discovered edaphochelin A, a previously unreported NRP that kills multi-drug-resistant Candida auris and Candida glabrata by disrupting respiratory chain proteins. Structural elucidation via nuclear magnetic resonance and mass spectrometry, alongside in vitro and in vivo validation, confirmed its efficacy, safety, and a mode of action distinct from existing antifungals-establishing edaphochelin A as a compelling drug candidate and NPDiscover as a powerful engine for scalable natural product discovery.

CP: biotechnology↗

Challenges and future directions in AI-driven biomaterials for microbiome-associated oral infectious diseases: A systematic review.

Oral biofilm-induced antimicrobial resistance is the core pathogenic mechanism of microbiome-associated oral infectious diseases (dental caries, periodontitis, peri-implantitis, and endodontic infection). Traditional therapies and biomaterials are limited by poor biofilm penetration, drug resistance induction, single functionality, and inadequate adaptation to dynamic oral microenvironmental changes (e.g., pH fluctuations, salivary rinsing, masticatory stimulation). Artificial intelligence (AI) has transformed the field by integrating materials science, microbiology, and stomatology data. Via machine learning, deep learning, and multi-physics simulation, AI optimizes biomaterial physicochemical properties, decodes microenvironmental signals, constructs precise sensing-response loops, and supports the full chain of material design, performance prediction, and action simulation, advancing treatment from empirical intervention to precision regulation. This systematic review retrieved literature from PubMed, Embase, and Web of Science (January 2016-January 2026) using keywords across three dimensions: AI, biomaterials, and oral microbiome. Following inclusion/exclusion criteria, 99 articles were included. It elaborates on five core mechanisms of AI-driven oral biomaterials (precise oral microbiome analysis, targeted material design/optimization, performance prediction/simulation, targeted delivery/intervention, effect evaluation/dynamic regulation), analyzes their applications in microbiome-targeted biomaterial research and development (R&D) and clinical practice for the four major oral infectious diseases, addresses technical bottlenecks (insufficient targeting specificity and precision of biomaterials, poor stability and durability in complex oral microenvironments, inadequate biofilm disruption capacity, and clinical translation obstacles), and proposes future directions (multimodal design to enhance targeting specificity, structural and component optimization to improve stability/durability, development of multi-mechanism synergistic biofilm disruption strategies, strengthening translational research for clinical application, and deep integration of AI in the full chain of biomaterial R&D). This work provides comprehensive theoretical and practical support for the R&D, optimization, and clinical translation of AI-driven microbiome-targeted oral biomaterials.

Humans↗

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95 % CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma↗

Association of long-term exposure to ambient air pollution and myopia in Chinese children.

Ambient air pollution is recognized as a major global health concern, but evidence on its association with childhood myopia remains limited, particularly under multi-pollutant exposure conditions. A school-based study was conducted in Tianjin, China, including 212,566 students in grades 4-6. The 3-yr mean concentrations of particulate matter with aerodynamic diameter ≤ 2.5 μm (PM2.5), its major components (sulfate (SO42-), nitrate (NO3-), ammonium (NH4+), organic matter (OM), and black carbon (BC)), and ozone (O3) were estimated using machine-learning exposure models and linked to school locations. Restricted cubic splines and quartile-based modified Poisson models were used to assess single-pollutant exposure-response relationships, and quantile-based g-computation was applied to estimate joint pollutant associations. In single-pollutant models, the highest quartile of SO42- was associated with higher myopia prevalence compared with the lowest quartile (PR = 1.10; 95% CI, 1.07-1.13). O3 showed weaker and non-monotonic positive patterns (Q4 vs Q1: PR = 1.03; 95% CI, 1.00-1.05). In mixture analyses, a one-quartile increase in joint exposure was associated with higher myopia prevalence (PR = 1.017; 95% CI, 1.007-1.027). Sensitivity analyses generally supported the direction of the main findings. These findings suggest that long-term exposure to specific ambient air pollutants may be associated with myopia in school-aged children.

Chemical components↗

Prediction of methylated CpGs in DNA sequences using a support vector machine.

DNA methylation plays a key role in the regulation of gene expression. The most common type of DNA modification consists of the methylation of cytosine in the CpG dinucleotide. At the present time, there is no method available for the prediction of DNA methylation sites. Therefore, in this study we have developed a support vector machine (SVM)-based method for the prediction of cytosine methylation in CpG dinucleotides. Initially a SVM module was developed from human data for the prediction of human-specific methylation sites. This module achieved a MCC and AUC of 0.501 and 0.814, respectively, when evaluated using a 5-fold cross-validation. The performance of this SVM-based module was better than the classifiers built using alternative machine learning and statistical algorithms including artificial neural networks, Bayesian statistics, and decision trees. Additional SVM modules were also developed based on mammalian- and vertebrate-specific methylation patterns. The SVM module based on human methylation patterns was used for genome-wide analysis of methylation sites. This analysis demonstrated that the percentage of methylated CpGs is higher in UTRs as compared to exonic and intronic regions of human genes. This method is available on line for public use under the name of Methylator at http://bio.dfci.harvard.edu/Methylator/.

Algorithms↗

Non-destructive prediction of lead content in oilseed rape leaves by fluorescence hyperspectral technology based on neural network.

Based on fluorescence hyperspectral imaging (FHSI), this study targeted rapid, non-destructive quantification of lead (Pb) content in oilseed rape leaves treated with varying silicon (Si) concentrations, acquiring fluorescence spectra over the 484.43-1001.61 nm wavelength range. To optimize spectral data quality, preprocessing methods (Savitzky-Golay smoothing, first derivative, detrending) were comprehensively compared. Characteristic wavelengths were then selected via interval variable iterative shrinkage, which effectively compressed data dimensionality and reduced computational load. A hybrid SE-CL1DA model, fusing a 1D convolutional neural network, a long short-term memory network and SE attention mechanism was constructed, with Bayesian optimization tuning hyperparameters to boost stability. The BO-SE-CL1DA outperformed both traditional machine learning and insufficiently optimized deep learning model (Rp2=0.9609, RMSE = 0.0377 mg/kg, RPD = 5.1736), thus enabling accurate Pb estimation, supporting Si-regulated heavy metal stress management and facilitating agricultural contamination monitoring.

Plant Leaves↗

Implementation of automated signal generation in pharmacovigilance using a knowledge-based approach.

Automated signal generation is a growing field in pharmacovigilance that relies on data mining of huge spontaneous reporting systems for detecting unknown adverse drug reactions (ADR). Previous implementations of quantitative techniques did not take into account issues related to the medical dictionary for regulatory activities (MedDRA) terminology used for coding ADRs. MedDRA is a first generation terminology lacking formal definitions; grouping of similar medical conditions is not accurate due to taxonomic limitations. Our objective was to build a data-mining tool that improves signal detection algorithms by performing terminological reasoning on MedDRA codes described with the DAML+OIL description logic. We propose the PharmaMiner tool that implements quantitative techniques based on underlying statistical and bayesian models. It is a JAVA application displaying results in tabular format and performing terminological reasoning with the Racer inference engine. The mean frequency of drug-adverse effect associations in the French database was 2.66. Subsumption reasoning based on MedDRA taxonomical hierarchy produced a mean number of occurrence of 2.92 versus 3.63 (p < 0.001) obtained with a combined technique using subsumption and approximate matching reasoning based on the ontological structure. Semantic integration of terminological systems with data mining methods is a promising technique for improving machine learning in medical databases.

Adverse Drug Reaction Reporting Systems↗

Flow rate of some pharmaceutical diluents through die-orifices relevant to mini-tableting.

The effects of cylindrical orifice length and diameter on the flow rate of three commonly used pharmaceutical direct compression diluents (lactose, dibasic calcium phosphate dihydrate and pregelatinised starch) were investigated, besides the powder particle characteristics (particle size, aspect ratio, roundness and convexity) and the packing properties (true, bulk and tapped density). Flow rate was determined for three different sieve fractions through a series of miniature tableting dies of different orifice diameter (0.4, 0.3 and 0.2 cm) and thickness (1.5, 1.0 and 0.5 cm). It was found that flow rate decreased with the increase of the orifice length for the small diameter (0.2 cm) but for the large diameter (0.4 cm) was increased with the orifice length (die thickness). Flow rate changes with the orifice length are attributed to the flow regime (transitional arch formation) and possible alterations in the position of the free flowing zone caused by pressure gradients arising from the flow of self-entrained air, both above the entrance in the die orifice and across it. Modelling by the conventional Jones-Pilpel non-linear equation and by two machine learning algorithms (lazy learning, LL, and feed-forward back-propagation, FBP) was applied and predictive performance of the fitted models was compared. It was found that both FBP and LL algorithms have significantly higher predictive performance than the Jones-Pilpel non-linear equation, because they account both dimensions of the cylindrical die opening (diameter and length). The automatic relevance determination for FBP revealed that orifice length is the third most influential variable after the orifice diameter and particle size, followed by the bulk density, the difference between bulk and tapped densities and the particle convexity.

Algorithms↗

Integrative multi-omics profiling deciphers tumor microenvironment heterogeneity and immunotherapy vulnerabilities in lung neuroendocrine carcinomas.

INTRODUCTION: Lung neuroendocrine carcinomas (Lu-NECs) are rare, highly aggressive lung tumors with poor prognosis and limited therapeutic options. Understanding the tumor immune microenvironment (TIME) is crucial towards personalized therapeutic strategies. OBJECTIVES: This study aims to systematically characterize the heterogeneity and complexity of the TIME in Lu-NECs by integrating proteomic, transcriptomic, and genomic data. METHODS: We performed comprehensive immune-proteomic profiling of 76 Lu-NECs across diverse histopathological subtypes to elucidate intra-tumoral TIME heterogeneity at the proteomic level. Validation was conducted in multiple independent cohorts, including 112 Lu-NECs using immunohistochemistry, 147 Lu-NECs, and 17 small cell lung carcinoma samples using transcriptomics. We integrated proteomic, transcriptomic, genomic, and clinical data to assess molecular, immunological, and clinical features, as well as therapeutic vulnerabilities across different immune subtypes. RESULTS: We delineated the immuno-proteomic landscape of Lu-NECs and identified two major immuno-proteomic clusters with distinct immunological, molecular, and clinical characteristics. IPC1 was characterized by high immune cell infiltration, while IPC2 exhibited sparse immune cell presence. Genomic analysis revealed distinct mutational patterns, with IPC1 showing a higher incidence of APOBEC-associated mutation signatures and IPC2 being enriched for mutations associated with defective DNA mismatch repair and tobacco-related mutagens. Functional analyses indicated that IPC1 was related to immune and oncogenic signaling activity, whereas IPC2 was associated with cancer stemness and proliferation-related features. Furthermore, IPC1 and IPC2 demonstrated histological subtype-specific clinical benefits from postoperative chemotherapy. Finally, we developed a machine learning model (iPROM) to predict Lu-NECs immune classification and improve risk stratification, which was validated across multiple independent cohorts. CONCLUSIONS: This study advances the understanding of the tumor immune microenvironment in Lu-NECs through multi-omics characterization and highlights potential personalized therapeutic vulnerabilities tailored to the specific immune landscapes of Lu-NECs.

Humans↗

Proteomic signature of dementia risk in type 2 diabetes.

INTRODUCTION: Type 2 diabetes (T2D) significantly increases dementia risk, yet the molecular mechanisms underlying this association remain unclear. OBJECTIVES: This study aimed to identify protein signatures that distinguish dementia risk in T2D patients, develop a proteomic prediction model, and elucidate biological pathways connecting T2D and dementia. METHODS: We analyzed 2,920 plasma proteins from 52,958 participants (including 3,292 with T2D) in the UK Biobank Pharma Proteomics Project with a median follow-up of 14.6&#xa0;years. Cox regression models with interaction terms identified T2D-specific protein associations with dementia risk. Machine learning models were developed to predict dementia in T2D patients. Pathway analysis and weighted gene co-expression network analysis identified biological mechanisms linking T2D and dementia. RESULTS: We identified 471 proteins with significant interaction effects between T2D and dementia risk. In non-T2D individuals, elevated levels of neuronal pentraxin receptor (NPTXR, HR&#xa0;=&#xa0;0.74, 95&#xa0;%CI:0.66-0.83) and carbonic anhydrase 14 (CA14, HR&#xa0;=&#xa0;0.67, 95&#xa0;%CI:0.60-0.75) were exclusively associated with decreased dementia risk. Conversely, in T2D patients, elevated rho guanine nucleotide exchange factor 12 (ARHGEF12, HR&#xa0;=&#xa0;1.45, 95&#xa0;%CI:1.10-1.91) was specifically associated with increased dementia risk. A 51-protein model accurately predicted 15-year dementia risk in T2D patients (AUC&#xa0;=&#xa0;0.835, C-index&#xa0;=&#xa0;0.829), outperforming conventional clinical risk scores and maintaining high accuracy for Alzheimer's disease and vascular dementia. Pathway analysis revealed enrichment of IL6-JAK-STAT3 signaling in T2D-related dementia, while dysregulation of fatty acid metabolism was specific to T2D-associated Alzheimer's disease. CONCLUSIONS: This large-scale proteomic analysis identifies specific molecular signatures that differentiate dementia risk in diabetic and non-diabetic populations, with potential applications for early risk stratification and targeted interventions. The identified pathways provide novel insights into the pathophysiological processes connecting T2D and dementia and suggest potential therapeutic targets.

Humans↗

SILVER helps assign peptides to tandem mass spectra using intensity-based scoring.

Tandem mass spectrometry is commonly used to identify peptides (and thereby proteins) that are present in complex mixtures. Peptide identification from tandem mass spectra is partially automated, but still requires human curation to resolve "borderline" peptide-spectrum matches (PSMs). SILVER is web-based software that assists manual curation of tandem mass spectra, using a recently developed intensity-based machine-learning approach to scoring PSMs, Elias et al. In this method, a large training set of peptide, fragment, and peak-intensity properties for both matched and mismatched PSMs was used to develop a score measuring consistency between each predicted fragment ion of a candidate peptide and its corresponding observed spectral peak intensity. The SILVER interface provides a visual representation of match quality between each candidate fragment ion and the observed spectrum, thereby expediting manual curation of tandem mass spectra. SILVER is available online at http://llama.med.harvard.edu/Software.html.

Amino Acid Sequence↗