Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Recent advances in computational prediction of drug absorption and permeability in drug discovery.

Approximately 40%-60% of developing drugs failed during the clinical trials because of ADME/Tox deficiencies. Virtual screening should not be restricted to optimize binding affinity and improve selectivity; and the pharmacokinetic properties should also be included as important filters in virtual screening. Here, the current development in theoretical models to predict drug absorption-related properties, such as intestinal absorption, Caco-2 permeability, and blood-brain partitioning are reviewed. The important physicochemical properties used in the prediction of drug absorption, and the relevance of predictive models in the evaluation of passive drug absorption are discussed. Recent developments in the prediction of drug absorption, especially with the application of new machine learning methods and newly developed software are also discussed. Future directions for research are outlined.

Computer Simulation↗

Making sense of molecular signatures in the immune system.

The development of Functional Genomics technologies has opened new avenues to investigate the complexity of the immune system. Microarray technology has been particularly successful because of its relatively low cost and high genome coverage. Consequently to our ability to monitor the expression of a significant proportion of an organism genome, our understanding of the molecular dynamics behind cell differentiation and cell response has greatly improved. Molecular signatures associated to immune cells have provided important tools to investigate the molecular basis of diseases and have been often associated to diagnostic and prognostic markers. The availability of such large collection of data has stimulated the application of complex machine learning techniques in the attempt to link molecular signatures and cell physiology. Here we review the most recent developments in the analysis of molecular signatures in the immune system.

Animals↗

Computational models for predicting interactions with cytochrome p450 enzyme.

Cytochrome p450 (CYP) enzymes are predominantly involved in Phase 1 metabolism of xenobiotics. As only 6 isoenzymes are responsible for approximately 90 % of known oxidative drug metabolism, a number of frequently prescribed drugs share the CYP-mediated metabolic pathways. Competing for a single enzyme by the co-administered therapeutic agents can substantially alter the plasma concentration and clearance of the agents. Furthermore, many drugs are known to inhibit certain p450 enzymes which they are not substrates for. Because some drug-drug interactions could cause serious adverse events leading to a costly failure of drug development, early detection of potential drug-drug interactions is highly desirable. The ultimate goal is to be able to predict the CYP specificity and the interactions for a novel compound from its chemical structure. Current computational modeling approaches, such as two-dimensional and three-dimensional quantitative structure-activity relationship (QSAR), pharmacophore mapping and machine learning methods have resulted in statistically valid predictions. Homology models have been often combined with 3D-QSAR models to impose additional steric restrictions and/or to identify the interaction site on the proteins. This article summarizes the available models, methods, and key findings for CYP1A2, 2A6, 2C9, 2D6 and 3A4 isoenzymes.

Computational Biology↗

Toward a Better Paradigm for Head and Neck Cancer Treatment Applying AI (HNC-TACTIC): Protocol for an International Cohort Study of Electronic Health Records.

BACKGROUND: Head and neck squamous cell carcinomas (HNSCCs) cause considerable morbidity and mortality. Multimodal treatment strategies can cause significant toxicity, and therapy options are limited for recurrent disease. Immunotherapy has emerged as a promising approach. However, patient response variability underscores the need for better predictive markers. OBJECTIVE: This study aims to use artificial intelligence to develop two predictive models in patients with HNSCC to assess (1) progression or recurrence following primary curative treatment and (2) long-term survival after immunotherapy schemes in recurrent and metastatic disease. This study will also describe the characteristics of patients with early, locally advanced, and recurrent or metastatic cancers. METHODS: This is a retrospective, observational study of data captured in electronic health records (EHRs) from participating hospitals between January 1, 2014, and December 31, 2021. This study's population comprises adults diagnosed with HNSCC at any stage. Study variables, including demographics, comorbidities, clinical variables, treatments, and outcomes, will be extracted using EHRead, a technology that applies natural language processing and machine learning to extract and analyze structured and unstructured clinical information in deidentified EHRs. Predictive models based on dynamic risk stratification for treatment response and progression or recurrence will be developed using multivariable logistic regressions, decision tree classifiers, and random forest approaches. Descriptive and outcome analyses will be shown for different anatomic subsites and stratified by stage and treatment. RESULTS: This study began enrolling sites in July 2021 and is currently ongoing. By December 2025, data from 10 centers has been collected, comprising a total of 151,934,990 EHRs from 2,159,719 patients. CONCLUSIONS: Development of predictive models using artificial intelligence will advance clinical understanding of HNSCC to improve patient outcomes.

Humans↗

Classifying gene expression profiles from pairwise mRNA comparisons.

We present a new approach to molecular classification based on mRNA comparisons. Our method, referred to as the top-scoring pair(s) (TSP) classifier, is motivated by current technical and practical limitations in using gene expression microarray data for class prediction, for example to detect disease, identify tumors or predict treatment response. Accurate statistical inference from such data is difficult due to the small number of observations, typically tens, relative to the large number of genes, typically thousands. Moreover, conventional methods from machine learning lead to decisions which are usually very difficult to interpret in simple or biologically meaningful terms. In contrast, the TSP classifier provides decision rules which i) involve very few genes and only relative expression values (e.g., comparing the mRNA counts within a single pair of genes); ii) are both accurate and transparent; and iii) provide specific hypotheses for follow-up studies. In particular, the TSP classifier achieves prediction rates with standard cancer data that are as high as those of previous studies which use considerably more genes and complex procedures. Finally, the TSP classifier is parameter-free, thus avoiding the type of over-fitting and inflated estimates of performance that result when all aspects of learning a predictor are not properly cross-validated.

Journal Article↗

Allostatic load is associated with symptoms in chronic fatigue syndrome patients.

OBJECTIVES: To further explore the relationship between chronic fatigue syndrome (CFS) and allostatic load (AL), we conducted a computational analysis involving 43 patients with CFS and 60 nonfatigued, healthy controls (NF) enrolled in a population-based case-control study in Wichita (KS, USA). We used traditional biostatistical methods to measure the association of high AL to standardized measures of physical and mental functioning, disability, fatigue and general symptom severity. We also used nonlinear regression technology embedded in machine learning algorithms to learn equations predicting various CFS symptoms based on the individual components of the allostatic load index (ALI). METHODS: An ALI was computed for all study participants using available laboratory and clinical data on metabolic, cardiovascular and hypothalamic-pituitary-adrenal (HPA) axis factors. Physical and mental functioning/impairment was measured using the Medical Outcomes Study 36-item Short Form Health Survey (SF-36); current fatigue was measured using the 20-item multidimensional fatigue inventory (MFI); frequency and intensity of symptoms was measured using the 19-item symptom inventory (SI). Genetic programming, a nonlinear regression technique, was used to learn an ensemble of different predictive equations rather just than a single one. Statistical analysis was based on the calculation of the percentage of equations in the ensemble that utilized each input variable, producing a measure of the 'utility' of the variable for the predictive problem at hand. Traditional biostatistics methods include the median and Wilcoxon tests for comparing the median levels of subscale scores obtained on the SF-36, the MFI and the SI summary score. RESULTS: Among CFS patients, but not controls, a high level of AL was significantly associated with lower median values (indicating worse health) of bodily pain, physical functioning and general symptom frequency/intensity. Using genetic programming, the ALI was determined to be a better predictor of these three health measures than any subcombination of ALI components among cases, but not controls.

Adult↗

Using data preprocessing and single layer perceptron to analyze laboratory data.

During daily work in hospitals a large amount of clinical data is produced each day. Totally computerized patient records are not yet widely used but a large part of essential information is already stored on computer files. These include laboratory test results, diagnoses, codes for operations, codes of histopathological diagnoses and maybe even the patient's medication. Accordingly, these databases include much clinical knowledge that would be useful for clinicians. Laboratories try to support clinicians by producing reference values for laboratory tests. It is, of course, necessary information but, however, it does not give very much information about the weight of evidence that an abnormal laboratory test will give in special clinical settings. We have developed a software package - DiagaiD - in order to build a smart link between patient databases and clinicians. It utilizes neural network-based machine learning techniques and can produce decision support which meets the special needs of clinicians. From example cases it can learn clinically relevant transformations from original numeric values to logical values. By using data transformation together with a single layer perceptron it is possible to build nonlinear models from a set of preclassified example cases. In this paper, we use two small datasets to show how this scheme works in the diagnosis of acute appendicitis and in the diagnosis of myocardial infarction. Results are compared with those obtained using logistic regression or backpropagation neural networks. The performance of our neuro-fuzzy tool seemed to be slightly better in these two materials but the differences did not reach statistical significance.

Appendicitis↗

A case-acquisition and decision-support system for the analysis of group-average lactation curves.

A case-acquisition and decision-support system was developed to support the analysis of group-average lactation curves and to acquire example cases from domain specialists. This software was developed through several iterations of a three-step approach involving 1) problem analysis and formulation in consultation with two dairy nutrition specialists; 2) development of a case-acquisition and decision-support prototype by the system developer; and 3) use of the prototype by the domain specialists to analyze and classify milk-recording data from example herds. The overall problem was decomposed into three subproblems: removal of outlier tests and lactation curves of individual cows; interpretation of group-average lactation curves; and diagnosis of detected abnormalities at the herd level through the identification of potential management deficiencies. For each subproblem, a software module was developed allowing the user to analyze both graphical and numerical performance representations and classify these representations using predefined linguistic descriptors. The example-based method for the development of the program proved to be very useful, facilitating the communication between system developer and domain specialists, and allowing the specialists to explore the appropriateness of the various prototypes developed. The resulting software represents a formalization of the approach to group-average lactation curve analysis, elicited from the two domain specialists. In future research, the case-acquisition and decision-support system will be complemented with knowledge to automate identified classification tasks, which will be captured through the application of machine-learning techniques to example cases, acquired from domain specialists using the software.

Animals↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

Advancing precision tacrolimus therapy: a systems genetics dissection in BXD platform.

BACKGROUND: Tacrolimus is a core immunosuppressant in organ transplantation, but its narrow therapeutic window and significant pharmacokinetic variability hinder precision dosing. Although CYP3A5-guided strategies have established clinical relevance for tacrolimus initial dose adjustment, they do not fully account for the marked interindividual variability in tacrolimus exposure, highlighting the need for complementary models to decode more complex genetic regulation. This study aimed to identify candidate genetic modulators of tacrolimus metabolism and develop an integrated predictive framework for individualized therapy. METHODS: Using 46 BXD recombinant inbred mouse strains, we characterized transcriptomics and machine learning, and validated key genes. We then constructed a clinical model using data from 168 renal transplant recipients. RESULTS: We identified 19 genomic loci associated with tacrolimus pharmacokinetic traits and supported DBP/CYP2A6 as candidate modulators associated with tacrolimus disposition. The clinical prediction model, incorporating these genes and clinical variables, achieved robust AUROC. CONCLUSIONS: These findings support a polygenic contribution to tacrolimus metabolism and provide an experimental and computational framework for identifying candidate modulators relevant to individualized dosing. The BXD mouse platform offers a systems-genetics approach for mechanistic discovery that may inform future translational studies on tacrolimus precision dosing.

Animals↗

Exposome influences: a multi-omics perspective on the combined toxic effects of pharmaceuticals and personal care products in Alzheimer's disease.

According to WHO data, approximately 57 million people worldwide were affected by dementia in 2021, with prevalence projected to rise. Alzheimer's disease (AD), responsible for 60%-80% of dementia cases, continues to be a leading cause of mortality, with current treatments offering limited efficacy and disease-modifying therapies lacking widespread adoption or conclusive safety evidence, shifting the focus toward prevention and risk modification. Risk factors for AD include both non-modifiable elements, such as age, genetics, and gender, and modifiable factors, like environmental pollution, health status, and diet. While age remains the primary non-modifiable risk factor, early-onset dementia represents only up to 9% of cases. Addressing modifiable factors is essential, as it could prevent or delay almost half of dementia cases, with interventions-such as increased physical activity, smoking cessation, alcohol limitation, and overall health management-being significantly associated with a reduced risk. In this context, the exposome approach offers a comprehensive, integrative framework in which both modifiable and non-modifiable risk factors interact to influence individual susceptibility. Within the neural exposome, chronic low-dose exposure to xenobiotics-such as industrial chemicals, pesticides, metals, pharmaceuticals and personal care products (PPCPs), and air pollutants-may induce neurodegeneration via mechanisms including oxidative stress, neuroinflammation, proteinopathies, and epigenetic modifications, although establishing causality remains challenging. Integration of genomics, transcriptomics, proteomics, metabolomics, and lipidomics, combined with artificial intelligence (AI) techniques such as machine learning (ML) and deep learning (DL), provides promising avenues for biomarker discovery, enhanced preventive strategies, early non-invasive diagnosis, and therapeutic target identification by integrating multi-layered biological data with exposure profiles. This review highlights emerging AD risk factors-including PPCPs-underscoring complex, multifactorial nature of AD and exposome, and the requirement for an interdisciplinary research approach, while also addressing several critical research gaps and methodological limitations.

Alzheimer’s disease↗

The Computational Revolution in Natural Product Research: A Data-Driven Roadmap for Next-Generation Drug Development.

Natural products (NPs) have historically provided the foundational scaffolds for drug development, yet traditional bioprospecting faces critical limitations: high rediscovery rates, laborious isolation workflows, and substantial attrition during clinical translation. The emergence of big data technologies is fundamentally transforming this landscape, enabling a shift from serendipity-based discovery toward systematic, data-driven approaches. This review examines how the integration of artificial intelligence (AI), machine learning (ML), and multi-omics datasets is accelerating natural product research across three key domains: (1) genome mining for biosynthetic gene cluster identification using platforms such as antiSMASH, (2) cheminformatics-driven prediction of structure-activity relationships and ADMET properties, and (3) metabolomics-guided dereplication to prioritize novel bioactive scaffolds. We evaluate the convergence of genomics, metabolomics, and computational chemistry in enabling in silico lead optimization and the discovery of cryptic metabolites from previously inaccessible microbial taxa. While challenges in data standardization and scalability persist, the synergy between big data and NP research is accelerating clinical translation. Despite persistent challenges in data standardization, scalability, and equitable benefit-sharing, the convergence of big data and NP research is poised to redefine drug development. These advances position computational NP research as a cornerstone of next-generation drug development.

big data analytics↗

Evaluation of unsupervised semantic mapping of natural language with Leximancer concept mapping.

The Leximancer system is a relatively new method for transforming lexical co-occurrence information from natural language into semantic patterns in a nunsupervised manner. It employs two stages of co-occurrence information extraction-semantic and relational-using a different algorithm for each stage. The algorithms used are statistical, but they employ nonlinear dynamics and machine learning. This article is an attempt to validate the output of Leximancer, using a set of evaluation criteria taken from content analysis that are appropriate for knowledge discovery tasks.

Equipment Design↗

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products↗

Placenta-derived Exosomes Mitigate Hypoxia-Induced Trophoblast Apoptosis and Inflammatory Progression via SASH1.

SASH1 is a signal adaptor protein involved in cell growth, apoptosis, and immune regulation, and has been increasingly studied in tumor and immune cells. Emerging evidence suggests that SASH1 plays an important role in inflammatory responses and cellular homeostasis, processes that are closely associated with the development of PE. This study aimed to determine whether SASH1 contributes to trophoblast apoptosis and inflammatory responses in PE and whether P-EXOS exerts protective effects through SASH1 regulation. In this study, three PE-related transcriptomic datasets (GSE75010, GSE10588, and GSE60438) were analyzed to identify shared differentially expressed genes (DEGs), followed by Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. Machine learning algorithms were further applied to screen key candidate genes, and single-cell RNA sequencing data were used to characterize cellular heterogeneity in placental tissue and to determine cell type-specific expression patterns. SASH1 was identified as a consensus candidate gene and was significantly upregulated in trophoblast cells from PE samples. In vitro, a hypoxia-treated HTR-8/SVneo trophoblast cell model was established, combined with SASH1 knockdown, SASH1 overexpression, and co-culture with P-EXOS. Functional experiments showed that knockdown of SASH1 significantly suppressed hypoxia-induced trophoblast apoptosis and reduced the secretion of pro-inflammatory cytokines, including IL-6, IL-1β, and TNF-α, whereas SASH1 overexpression promoted apoptosis and inflammatory responses. In addition, P-EXOS treatment markedly reduced SASH1 expression at both mRNA and protein levels and attenuated hypoxia-induced trophoblast injury, while SASH1 overexpression largely abolished these protective effects. Taken together, these findings indicate that SASH1 plays a critical role in trophoblast apoptosis and inflammatory responses in PE. P-EXOS may alleviate hypoxia-induced trophoblastic injury by suppressing SASH1 expression, providing new insights into the molecular mechanisms and potential therapeutic targets for PE.

Trophoblasts↗

LINC01871-Mediated Sensitivity to Cyclin-Dependent Kinase 4/6 Inhibitors in Human Breast Cancer.

Breast cancer remains the most frequently diagnosed malignancy in women, and resistance to cyclin-dependent kinase 4 and 6 (CDK4/6) inhibitors limits long-term treatment efficacy. This study aimed to identify long non-coding RNAs (lncRNAs) associated with predicted sensitivity to CDK4/6 inhibitors and to investigate their biological functions in breast cancer. Transcriptomic data from The Cancer Genome Atlas (TCGA) and drug sensitivity data from the Genomics of Drug Sensitivity in Cancer 2 (GDSC2) database were integrated, and drug sensitivity was predicted using the oncoPredict algorithm. Candidate lncRNAs were identified through differential expression analysis, weighted gene co-expression network analysis, prognostic analysis, and machine learning. The biological functions of LINC01871 were subsequently evaluated using in vitro and in vivo experiments. Sixty-two lncRNAs associated with predicted sensitivity to ribociclib and palbociclib were identified, and six core lncRNAs were selected. LINC01871 showed the highest discriminatory performance for predicted drug sensitivity. Overexpression of LINC01871 was associated with increased sensitivity of breast cancer cells to ribociclib and palbociclib, inhibition of cell proliferation, promotion of apoptosis, and suppression of nuclear factor kappa B (NF-κB) signaling. Single-cell transcriptomic analysis demonstrated high LINC01871 expression in T cells and natural killer (NK) cells, while transcriptome-based immune infiltration analyses showed that high LINC01871 expression was associated with increased immune infiltration. These findings identify LINC01871 as a candidate biomarker of sensitivity to CDK4/6 inhibitors and demonstrate its tumor-suppressive effects in breast cancer. Further clinical and mechanistic studies are required to validate its predictive value and therapeutic relevance.

Humans↗

Automated identification of cancerous smears using various competitive intelligent techniques.

In this study the performance of various intelligent methodologies is compared in the task of pap-smear diagnosis. The selected intelligent methodologies are briefly described and explained, and then, the acquired results are presented and discussed for their comprehensibility and usefulness to medical staff, either for fault diagnosis tasks, or for the construction of automated computer-assisted classification of smears. The intelligent methodologies used for the construction of pap-smear classifiers, are different clustering approaches, feature selection, neuro-fuzzy systems, inductive machine learning, genetic programming, and second order neural networks. Acquired results reveal the power of most intelligent techniques to obtain high quality solutions in this difficult problem of medical diagnosis. Some of the methods obtain almost perfect diagnostic accuracy in test data, but the outcome lacks comprehensibility. On the other hand, results scoring high in terms of comprehensibility are acquired from some methods, but with the drawback of achieving lower diagnostic accuracy. The experimental data used in this study were collected at a previous stage, for the purpose of combining intelligent diagnostic methodologies with other existing computer imaging technologies towards the construction of an automated smear cell classification device.

Artificial Intelligence↗

Using data mining to explore complex clinical decisions: A study of hospitalization after a suicide attempt.

BACKGROUND: Medical education is moving toward developing guidelines using the evidence-based approach; however, controlled data are missing for answering complex treatment decisions such as those made during suicide attempts. A new set of statistical techniques called data mining (or machine learning) is being used by different industries to explore complex databases and can be used to explore large clinical databases. METHOD: The study goal was to reanalyze, using data mining techniques, a published study of which variables predicted psychiatrists' decisions to hospitalize in 509 suicide attempters over the age of 18 years who were assessed in the emergency department. Patients were recruited for the study between 1996 and 1998. Traditional multivariate statistics were compared with data mining techniques to determine variables predicting hospitalization. RESULTS: Five analyses done by psychiatric researchers using traditional statistical techniques classified 72% to 88% of patients correctly. The model developed by researchers with no psychiatric knowledge and employing data mining techniques used 5 variables (drug consumption during the attempt, relief that the attempt was not effective, lack of family support, being a housewife, and family history of suicide attempts) and classified 99% of patients correctly (99% sensitivity and 100% specificity). CONCLUSIONS: This reanalysis of a published study fundamentally tries to make the point that these new multivariate techniques, called data mining, can be used to study large clinical databases in psychiatry. Data mining techniques may be used to explore important treatment questions and outcomes in large clinical databases and to help develop guidelines for problems where controlled data are difficult to obtain. New opportunities for good clinical research may be developed by using data mining analyses.

Adult↗