Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

SMIREP: predicting chemical activity from SMILES.

Most approaches to structure-activity-relationship (SAR) prediction proceed in two steps. In the first step, a typically large set of fingerprints, or fragments of interest, is constructed (either by hand or by some recent data mining techniques). In the second step, machine learning techniques are applied to obtain a predictive model. The result is often not only a highly accurate but also hard to interpret model. In this paper, we demonstrate the capabilities of a novel SAR algorithm, SMIREP, which tightly integrates the fragment and model generation steps and which yields simple models in the form of a small set of IF-THEN rules. These rules contain SMILES fragments, which are easy to understand to the computational chemist. SMIREP combines ideas from the well-known IREP rule learner with a novel fragmentation algorithm for SMILES strings. SMIREP has been evaluated on three problems: the prediction of binding activities for the estrogen receptor (Environmental Protection Agency's (EPA's) Distributed Structure-Searchable Toxicity (DSSTox) National Center for Toxicological Research estrogen receptor (NCTRER) Database), the prediction of mutagenicity using the carcinogenic potency database (CPDB), and the prediction of biodegradability on a subset of the Environmental Fate Database (EFDB). In these applications, SMIREP has the advantage of producing easily interpretable rules while having predictive accuracies that are comparable to those of alternative state-of-the-art techniques.

Algorithms↗

Evaluating the C-section rate of different physician practices: using machine learning to model standard practice.

The C-section rate of a population of 22,175 expectant mothers is 16.8%; yet the 17 physician groups that serve this population have vastly different group C-section rates, ranging from 13% to 23%. Our goal is to determine retrospectively if the variations in the observed rates can be attributed to variations in the intrinsic risk of the patient sub-populations (i.e. some groups contain more "high-risk C-section" patients), or differences in physician practice (i.e. some groups do more C-sections). We apply machine learning to this problem by training models to predict standard practice from retrospective data. We then use the models of standard practice to evaluate the C-section rate of each physician practice. Our results indicate that although there is variation in intrinsic risk among the groups, there also is much variation in physician practice.

Artificial Intelligence↗

Integrative machine learning models to unravel gut microbial dysbiosis and functional disruption in polycystic ovary syndrome.

OBJECTIVE: To study gut microbial diversity and metabolic pathway disruptions in women with PolyCystic Ovary Syndrome (PCOS) compared with healthy controls, and to evaluate the diagnostic potential of microbiome-driven machine learning models. DESIGN: Case-controlled metagenomic data analysis SUBJECTS: Gut metagenomic data from women diagnosed with PCOS and age-matched healthy female controls EXPOSURE: Presence of PCOS MAIN OUTCOME MEASURES: The primary outcome measures will include gut microbial alpha and beta diversity indices, microbial taxon abundance, functional pathway profiles, predicted metabolite levels, microbe-functional pathway-metabolite interaction networks, and the diagnostic accuracy of microbiome-based machine learning models. RESULTS: Alpha and beta diversity analyses revealed marked gut microbial dysbiosis in women with PCOS, despite comparable species richness to healthy controls. Differential abundance analysis identified 41 significantly altered microbial species, including enrichment of proinflammatory taxa, such as Bacteroides vulgatus and Ruminococcus gnavus, and depletion of beneficial commensals, including Roseburia hominis and Prevotella copri. These compositional shifts indicate a proinflammatory microbial community structure in PCOS. Functional profiling demonstrated the upregulation of pathways involved in nucleotide turnover, lipid and carbohydrate metabolism, and neurotransmitter synthesis, potentially contributing to metabolic and neuroendocrine disruption. Network analysis revealed fragmented and unstable microbial-metabolite associations in PCOS compared with cohesive networks in controls. Microbiome-based machine learning models achieved a diagnostic accuracy of 84.25% (area under the curve 0.93), underscoring their predictive potential. CONCLUSION: The gut microbiome in PCOS is characterized by a proinflammatory community structure and disrupted metabolic pathways. These findings demonstrate the diagnostic potential of microbiome-based models and underscore the gut microbiome as a promising target for therapeutic interventions in the management of PCOS.

Polycystic Ovary Syndrome↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 688 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

DNA methylation↗

Possible prediction of chemoradiosensitivity of esophageal cancer by serum protein profiling.

PURPOSE: Establishment of a reliable method of predicting the efficacy of chemotherapy and radiotherapy is necessary to provide the most suitable treatment for each cancer patient. We investigated whether proteomic profiles of serum samples obtained from untreated patients were capable of being used to predict the efficacy of combined preoperative chemoradiotherapy against esophageal cancer. EXPERIMENTAL DESIGN: Proteomic spectra were obtained from a training set of 27 serum samples (15 pathologically diagnosed responders to preoperative chemoradiotherapy and 12 nonresponders) by surface-enhanced laser desorption and ionization coupled with hybrid quadrupole time-of-flight mass spectrometry. A proteomic pattern prediction model was constructed from the training set by machine learning algorithms, and it was then tested with an independent validation set consisting of serum samples from 15 esophageal cancer patients in a blinded manner. RESULTS: We selected a set of four mass peaks, at 7,420, 9,112, 17,123, and 12,867 m/z, from a total of 859 protein peaks, as perfectly distinguishing responders from nonresponders in the training set with a support vector machine algorithm. This set of peaks (i.e., the classifier) correctly diagnosed chemoradiosensitivity in 93.3% (14 of 15) of the cases in the validation set. CONCLUSIONS: Recent mass spectrometric approaches have revealed that serum contains a large volume of information that reflects the microenvironment of diseased organs. Although a multi-institutional large-scale study will be necessary to confirm each component of the classifier, there is a subtle but definite difference in serum proteomic profile between responders and nonresponders to chemoradiotherapy.

Aged↗

Modelling the structure and function of enzymes by machine learning.

A machine learning program, GOLEM, has been applied to two problems: (1) the prediction of protein secondary structure from sequence and (2) modelling a quantitative structure-activity relationship in drug design. GOLEM takes as input observations and combines them with background knowledge of chemistry to yield rules expressed as stereochemical principles for prediction. The secondary structure prediction was explored on the alpha/alpha class of proteins; on an unrelated test set it yielded 81% accuracy. The rules from GOLEM defined patterns of residues forming alpha-helices. The system studied for drug design was the activities of trimethoprim analogues binding to E. coli dihydrofolate reductase. The GOLEM rules were a better model than standard regression approaches. More importantly, these rules described the chemical properties of the enzyme-binding site that were in broad agreement with the crystallographic structure.

Amino Acid Sequence↗

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans↗

MULTIPREVENT: Integrated screening for smoking-related multimorbidity using low-dose chest computed tomography.

OBJECTIVES: Tobacco consumption, combined with individual genetic predispositions, contributes to an age-dependent risk not only for lung cancer but also for other non-communicable diseases (NCDs) such as cardiovascular disease (CVD), chronic obstructive pulmonary disease (COPD), osteoporosis, and diabetes. The MULTIPREVENT project aims to validate whether low-dose computed tomography (LDCT) of the chest, combined with simple biomarkers, functional tests, and genomic profiling, can serve as an effective tool for comprehensive health assessment and risk prediction of multimorbidity in adults. STUDY DESIGN: The study is based on a prospective epidemiological design involving 3000 participants from the MOLTEST-BIS lung cancer screening cohort (2016-2018). These participants, aged 50-79 years (during MOLTEST-BIS) and with a smoking history of at least 30 pack-years, will undergo two follow-up assessments in 2025-2027 and 2030-2032. METHODS: Each follow-up includes LDCT, spirometry, standardized blood pressure measurement, anthropometric evaluation, biomarker assessment (lipid profile, lipoprotein(a), glycated haemoglobin), and health-related questionnaires. Genetic profiling will be performed using the Illumina Infinium Global Screening Arrays approach to identify inherited predispositions to major NCDs. All data, clinical, imaging (including radiomics), molecular, and genetic, will be integrated through machine learning algorithms to develop AI-based risk prediction models. RESULTS: The MULTIPREVENT study is expected to generate a wide range of scientific, clinical, and infrastructural results that will serve as a foundation for future public health initiatives in integrated prevention. CONCLUSIONS: By linking imaging and biochemical markers, genetic susceptibility, and clinical parameters within a longitudinal design, MULTIPREVENT will establish data-driven, AI-supported prevention strategies aimed at reducing morbidity and mortality among adults exposed to tobacco. The project will also serve as a model for population-based multimorbidity prevention programs.

Humans↗

The predictive and explanatory power of inductive decision trees: a comparison with artificial neural network learning as applied to the noninvasive diagnosis of coronary artery disease.

BACKGROUND: This paper compares two machine learning systems, an inductive decision tree (IDT) and a back-propagation neural network (ANN), in the noninvasive assessment of coronary artery disease given a set of diagnostic input attributes. A collection of 490 patient cases were accumulated from the reference of diagnostic stress myocardial scintigraphy performed in a nuclear medicine department. All cases had correlating angiography, the results of which were used to derive the target diagnoses. Input attributes included 4 baseline clinical characteristics, 4 nonimaging stress components, and 3 scintigraphic findings. METHODS: We chose 4 possible angiographic criteria for coronary artery disease and assessed the ability of each learning system to develop a diagnostic model. The 2 machine learning systems were compared on the basis of predictive performance and explanatory power. RESULTS: Cross-validation experiments showed the 2 machine learning systems to have equivalent predictive power at the same level as the clinical scan reading. For the 70% stenosis criterion, the IDT had a sensitivity of 94 +/- 3% (mean +/- 95% confidence interval) and a specificity of 59 +/- 8%, and the ANN had a sensitivity of 97 +/- 2% and a specificity of 51 +/- 13%. However the IDT system exhibited excellent explanatory power; producing simple representations of the diagnostic models which agree with previous research. CONCLUSION: In comparison with the more widely used ANNs, the IDT learning system may bring advantages to certain problems in diagnostic classification.

Coronary Angiography↗

The Path-A metabolic pathway prediction web server.

Pathway Analyst (Path-A) is a publicly available web server (http://path-a.cs.ualberta.ca) that predicts metabolic pathways. It takes a FASTA format file containing a set of query protein sequences from a single organism (a partial or complete proteome) and identifies those sequences that are likely to participate in any of its supported metabolic pathways (currently 10). Path-A uses a number of machine-learning and sequence analysis techniques (e.g. SVM, BLAST and HMM) to predict pathways. Each machine-learned classifier exploits similarity between sequences in the pathways of its model organisms and sequences in the query set. It predicts the pathways that are present in the query organism and annotates each predicted reaction and catalyst, using the appropriate sequences from the query set. Path-A also provides a browsable and searchable database of the pathways for the model organisms that are used to make its predictions. Path-A's predictor sets (using different classifier technologies) have been evaluated using standard cross-validation techniques on a dataset of 10 metabolic pathways across 13 model organisms--a total of 125 organism-specific pathways. The most accurate classifier technology obtained a mean precision of 78.3% and a mean recall of 92.6% in predicting all catalyst proteins, of all reactions, in all pathways present in the dataset. Although Path-A currently only supports metabolic pathways, the underlying prediction techniques are general enough for other types of pathways. Consequently, it is our intent to extend Path-A to predict other types of pathways, including signalling pathways.

Algorithms↗

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal↗

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase↗

Artificial Intelligence and Machine Learning Applications in Fibromuscular Dysplasia: Transforming Diagnosis, Risk Stratification, and Clinical Decision-Making.

Fibromuscular dysplasia (FMD) is a non-atherosclerotic vascular disorder with heterogeneous presentations, making diagnosis and management highly dependent on imaging and clinical expertise. This narrative review examines how artificial intelligence (AI) and machine learning (ML) are transforming FMD care. AI-enhanced imaging, particularly convolutional neural network-based analysis, improves detection of the characteristic "string-of-beads" pattern on CT angiography, magnetic resonance angiography, and ultrasound, although FMD-specific validation remains limited. ML models facilitate risk stratification, prediction of disease progression, and early identification of complications such as aneurysms and stroke by integrating clinical, imaging, and genomic data. AI-driven clinical decision support systems further enable personalized treatment selection through pharmacogenomic insights and robot-assisted interventions. Despite promising real-world applications, challenges persist, including limited large-scale datasets, workflow integration, regulatory barriers, and algorithmic bias affecting underrepresented populations. Future advances in explainable AI, federated learning, and digital health integration may enable a shift toward predictive, patient-centered FMD management.

Humans↗

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models↗

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

Predicting host tropism in influenza a viruses: insights from multi-segment nucleotide signatures.

BACKGROUND: Influenza A virus (IAV) poses a significant public health threat due to its cross-species transmission and complex host adaptation mechanisms. This study integrated whole-genome data from avian, human, swine, and bovine IAV strains, using machine learning to predict viral host tropism based on nucleotide site features and to identify key sites driving host adaptation along with their synergistic effects. METHODS: A total of 64,000 IAV sequences from avian, human, swine, and bovine hosts were analyzed to build host-prediction models. A four-class classification framework (avian, human, swine, bovine) was constructed using nucleotide site features from all eight genomic segments (PB2, PB1, PA, HA, NP, NA, MP, NS). Eight machine learning algorithms (logistic regression, decision tree, random forest, SVM, KNN, gradient boosting, XGBoost, LightGBM) were benchmarked via 10-fold stratified cross-validation. Model performance was evaluated using accuracy, precision, recall, F1-score, AUPRC, and AUC. SHAP (SHapley Additive exPlanations) analysis prioritized critical nucleotide sites, while bivariate association tests identified synergistic/antagonistic interactions between sites. Nucleotide composition profiles were compared across host groups using hierarchical clustering and heatmap visualization. RESULTS: The XGBoost algorithm demonstrated the best and most stable performance, achieving an AUC value of over 0.95 in distinguishing human-derived sequences from non-human ones. SHAP analysis identified the top 20 critical nucleotide sites for each gene segment, such as sites 46 and 698 in the NS segment. Nucleotide composition analysis revealed high similarity between human and swine sequences in the HA and PB2 segments, and between avian and bovine sequences. The HA segment was particularly challenging in differentiating human from swine strains. Bivariate site association analysis uncovered significant synergistic or antagonistic effects between key sites within gene segments, forming complex networks. For instance, in the NS segment, a positive prediction contribution was observed when sites 371, 698, and 419 were all G. CONCLUSIONS: This study advances our mechanistic understanding of IAV host adaptation, identifies molecular determinants for zoonotic risk stratification, and establishes a scalable machine learning framework for predicting viral host tropism through nucleotide signature analysis, thereby enhancing surveillance strategies and informing preventive measures against emerging viral threats.

Influenza A virus↗

Assessment of hepatotoxic liabilities by transcript profiling.

Male Wistar rats were treated with various model compounds or the appropriate vehicle controls in order to create a reference database for toxicogenomics assessment of novel compounds. Hepatotoxic compounds in the database were either known hepatotoxicants or showed hepatotoxicity during preclinical testing. Histopathology and clinical chemistry data were used to anchor the transcript profiles to an established endpoint (steatosis, cholestasis, direct acting, peroxisomal proliferation or nontoxic/control). These reference data were analyzed using a supervised learning method (support vector machines, SVM) to generate classification rules. This predictive model was subsequently used to assess compounds with regard to a potential hepatotoxic liability. A steatotic and a non-hepatotoxic 5HT(6) receptor antagonist compound from the same series were successfully discriminated by this toxicogenomics model. Additionally, an example is shown where a hepatotoxic liability was correctly recognized in the absence of pathological findings. In vitro experiments and a dog study confirmed the correctness of the toxicogenomics alert. Another interesting observation was that transcript profiles indicate toxicologically relevant changes at an earlier timepoint than routinely used methods. Together, these results support the useful application of toxicogenomics in raising alerts for adverse effects and generating mechanistic hypotheses that can be followed up by confirmatory experiments.

Animals↗

Estimating the association of antimicrobial resistance genes with minimum inhibitory concentration in Escherichia coli: an observational study.

BACKGROUND: Surveillance and prediction of antibiotic resistance in Escherichia coli relies on curated databases of genes and mutations. We aimed to quantify the effect of acquiring specific genetic elements on minimum inhibitory concentrations (MICs) for particular antibiotic-species combinations, addressing the current scarcity of such data in existing databases. METHODS: For this observational study, we evaluated a collection of E coli isolates with linked whole-genome sequencing and MIC data, originating from human urinary or bloodstream infections obtained from the Oxford University Hospitals National Health Service Foundation Trust in Oxfordshire, UK. We used multivariable interval regression models to estimate the change in MIC (with 95% CIs) for specific antibiotics associated with the acquisition of antibiotic resistance genes and associated mutations in the National Center for Biotechnology Information AMRFinder database, with and without an adjustment for population structure. We then tested the ability of these models to predict MIC and binary resistance or susceptibility using leave-one-out cross-validation. FINDINGS: We evaluated 2875 E coli isolates obtained during 2013-2018 and 2020. Although most ARGs and resistance mutations (89 [80%] of 111) were associated with an increased MIC, a much smaller number (27 [24%] of 111) was found to be putatively independently resistance-conferring (ie, associated with an MIC above the European Committee on Antimicrobial Susceptibility Testing breakpoint) when acquired in isolation. We found evidence of differential effects of acquired ARGs and resistance mutations between different generations of cephalosporin antibiotics and showed that sub-breakpoint variation in MIC can be linked to genetic mechanisms of resistance. 20 697 (83·3%; range 52·9-97·7 across all antibiotics) of 24 858 MICs were correctly exactly predicted and 23 677 (95·2%; 87·3-97·7) of 24 858 MICs were predicted to within one doubling dilution. INTERPRETATION: Quantitative estimates of the independent effect of the acquisition of ARGs on MIC add to the interpretability and utility of existing databases. Compared with approaches using machine learning models, the use of these estimates yields similar or better performance in the prediction of antibiotic resistance phenotype with more readily interpretable results. The methods outlined here could be readily applied to other antibiotic-pathogen combinations. FUNDING: The National Institute for Health and Care Research (NIHR) and the Medical Research Council (MRC).

Escherichia coli↗