Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Modeling obesity using abductive networks.

This paper investigates the use of abductive-network machine learning for modeling and predicting outcome parameters in terms of input parameters in medical survey data. Here we consider modeling obesity as represented by the waist-to-hip ratio (WHR) risk factor to investigate the influence of various parameters. The same approach would be useful in predicting values of clinical parameters that are difficult or expensive to measure from others that are more readily available. The AIM abductive network machine learning tool was used to model the WHR from 13 other health parameters. Survey data were collected for a randomly selected sample of 1100 persons aged 20 yr and over attending nine primary health care centers at Al-Khobar, Saudi Arabia. Models were synthesized by training on a randomly selected set of 800 cases, using both continuous and categorical representations of the parameters, and evaluated by predicting the WHR value for the remaining 300 cases. Models for WHR as a continuous variable predict the actual values within an error of 7.5% at the 90% confidence limits. Categorical models predict the correct logical value of WHR with an error in only 2 of the 300 evaluation cases. Analytical relationships derived from simple categorical models explain global observations on the total survey population to an accuracy as high as 99%. Simple continuous models represented as analytical functions highlight global relationships and trends. Results confirm the strong correlation between WHR and diastolic blood pressure, cholesterol level, and family history of obesity. Compared to other statistical and neural network approaches, AIM abductive networks provide faster and more automated model synthesis. A review is given of other areas where the proposed modeling approach can be useful in clinical practice.

Adult↗

Physicochemical characteristics of non-electrolytes and their uptake by Brugia pahangi and Dipetalonema viteae.

The uptake of a diverse set of 14C-labelled non-electrolytes by Brugia pahangi and Dipetalonema viteae was measured relative to the free diffusion of tritiated water. Inulin was used as a non-absorbable surface marker to account for non-electrolyte adherent to the surface of the parasite which had not crossed the cuticle. B. pahangi and D. viteae took up the non-electrolytes to a similar degree; a comparison of tissue uptake indices gave a correlation coefficient of 0.99. Worm uptake could not be described by non-electrolyte octanol/aqueous partition coefficients alone. However, greater success was achieved using further descriptors and pattern recognition techniques for data analysis. The whole molecule descriptors log P, molar refraction, melting point, dipole moment and CNDO total energy were obtained from computer chemistry and the literature. Using a linear learning machine to relate uptake to these 5 physicochemical descriptors it was possible to successfully classify non-electrolytes as high or low uptake. Multivariate regression analysis of uptake versus these 5 parameters gave a correlation coefficient of 0.77. However, this was not statistically significant and therefore could not be used for quantitative predictions of substance uptake by worms. This illustrates the value of 'pattern recognition' techniques such as the linear learning machine. Using such 'pattern recognition' methods on a chemically related set of compounds it is anticipated that predictions of uptake can be achieved and improved upon. Such predictions could then be used in drug design.

Animals↗

Firemaster 550 differentially alters gene expression underlying synaptic function in amygdala of prairie voles after gestational or lactational exposure.

Neurodevelopmental disorders often share similar behavioral diagnostic criteria including socioemotional and cognitive deficits. The prairie vole is a uniquely suitable model to study these deficits because they demonstrate strong social affiliation, bi-parental care, and partner attachment. Previously, we have shown that developmental exposure to the flame-retardant mixture Firemaster 550 (FM 550) impairs socioemotional behavior in the prairie vole and alters underlying neuroanatomy and function. However, the mechanisms for impaired pair bonding in males and increased anxiety in females remain unknown, along with the specific critical window(s) of vulnerability. Herein, we exposed prairie vole dams to FM 550 during gestation or lactation, and performed bulk RNA-seq on the amygdala, a hub of socioemotional processing, in their adult offspring. Two mathematically orthogonal methods were utilized for analysis, a linear statistical method and an ensemble machine learning method, incorporating sex as a biological variable. Gene ontology (GO) pathway analysis was performed following both and results compared to identify potential mechanisms of toxicity. GO results indicated consistent expression changes in the Synapse cellular component in all conditions, and implicated glutamatergic signaling specifically. Additionally, gestational exposure (GE) altered genes underlying modulation of synaptic transmission and neural development, while lactational exposure (LE) impacted genes underlying synaptic plasticity, axon guidance, and mitophagy. Machine learning identified disruption of endocrine system development, regulation of biosynthetic processes in GE animals, and suppression of various neuroinflammatory genes across multiple groups. Finally, we performed RNA expression analysis using Nanostring and demonstrated stronger correlation with the differentially expressed genes (DEG) of interest in females than males. Overall, this study demonstrates both the intersecting and distinct impacts of FM 550 exposure on amygdalar gene expression depending on sex and timing of exposure.

Animals↗

A neural-network based method for prediction of gamma-turns in proteins from multiple sequence alignment.

In the present study, an attempt has been made to develop a method for predicting gamma-turns in proteins. First, we have implemented the commonly used statistical and machine-learning techniques in the field of protein structure prediction, for the prediction of gamma-turns. All the methods have been trained and tested on a set of 320 nonhomologous protein chains by a fivefold cross-validation technique. It has been observed that the performance of all methods is very poor, having a Matthew's Correlation Coefficient (MCC) </= 0.06. Second, predicted secondary structure obtained from PSIPRED is used in gamma-turn prediction. It has been found that machine-learning methods outperform statistical methods and achieve an MCC of 0.11 when secondary structure information is used. The performance of gamma-turn prediction is further improved when multiple sequence alignment is used as the input instead of a single sequence. Based on this study, we have developed a method, GammaPred, for gamma-turn prediction (MCC = 0.17). The GammaPred is a neural-network-based method, which predicts gamma-turns in two steps. In the first step, a sequence-to-structure network is used to predict the gamma-turns from multiple alignment of protein sequence. In the second step, it uses a structure-to-structure network in which input consists of predicted gamma-turns obtained from the first step and predicted secondary structure obtained from PSIPRED.

Databases, Protein↗

Comparison of genetic algorithms and other classification methods in the diagnosis of female urinary incontinence.

Galactica, a newly developed machine-learning system that utilizes a genetic algorithm for learning, was compared with discriminant analysis, logistic regression, k-means cluster analysis, a C4.5 decision-tree generator and a random bit climber hill-climbing algorithm. The methods were evaluated in the diagnosis of female urinary incontinence in terms of prediction accuracy of classifiers, on the basis of patient data. The best methods were discriminant analysis, logistic regression, C4.5 and Galactica. Practically no statistically significant differences existed between the prediction accuracy of these classification methods. We consider that machine-learning systems C4.5 and Galactica are preferable for automatic construction of medical decision aids, because they can cope with missing data values directly and can present a classifier in a comprehensible form. Galactica performed nearly as well as C4.5. The results are in agreement with the results of earlier research, indicating that genetic algorithms are a competitive method for constructing classifiers from medical data.

Algorithms↗

Integrated multi-omics analyses identify an RAS-SLC11A2-associated molecular framework linking iron metabolism with PCOS-related cardiometabolic risk.

INTRODUCTION: PCOS is a common endocrine disorder with elevated cardiometabolic risk, yet the role of the renin-angiotensin system (RAS)-iron metabolism axis in this comorbidity remains unclear. We explored its underlying mechanisms and evaluated the therapeutic potential of gentiopicroside. METHODS: Integrated multi-omics analyses combining transcriptomics, single-cell RNA sequencing, Mendelian randomization, machine learning, molecular docking, and in vitro functional assays were performed to identify shared molecular pathways and therapeutic targets across PCOS, hypertension, NAFLD, and T2DM. RESULTS: SLC11A2 was consistently dysregulated in PCOS transcriptomic datasets, and associated with iron metabolism, inflammatory response and oxidative stress pathways. Genetic analyses validated RAS-related regulation in hypertension susceptibility and revealed shared genetic architecture between PCOS and cardiometabolic traits. Network and single-cell analyses characterized SLC11A2-associated molecular patterns in disease-relevant cell types; machine learning identified disease-classifying molecular signatures. Gentiopicroside alleviated inflammatory and oxidative stress phenotypes, including reduced IL-6 expression and reactive oxygen species accumulation. CONCLUSION: This study defines an RAS-SLC11A2 molecular framework linking iron metabolism dysregulation to PCOS-related cardiometabolic risk, elucidating the mechanisms connecting ovarian dysfunction, inflammation, oxidative stress and hypertension, and supports gentiopicroside as a promising therapeutic candidate.

Humans↗

Evaluation of automatic knowledge acquisition techniques in the diagnosis of acute abdominal pain. Acute Abdominal Pain Study Group.

Clinical diagnosis in acute abdominal pain is still a major problem. Computer-aided diagnosis offers some help; however, existing systems still produce high error rates. We therefore tested machine learning techniques in order to improve standard statistical systems. The investigation was based on a prospective clinical database with 1254 cases, 46 diagnostic parameters and 15 diagnoses. Independence Bayes and the automatic rule induction techniques ID3, NewId, PRISM, CN2, C4.5 and ITRULE were trained with 839 cases and separately tested on 415 cases. No major differences in overall accuracy were observed (43-48%), except for NewId, which was below the average. Between the different techniques some similarities were found, but also considerable differences with respect to specific diagnoses. Machine learning techniques did not improve the results of the standard model Independence Bayes. Problem dimensionality, sample size and model complexity are major factors influencing diagnostic accuracy in computer-aided diagnosis of acute abdominal pain.

Abdominal Pain↗

Comparing syntactic complexity in medical and non-medical corpora.

With the growing use of Natural Language Processing (NLP) techniques as solutions in Medical Informatics, the need to quickly and efficiently create the knowledge structures used by these systems has grown concurrently. Automatic discovery of a lexicon for use by an NLP system through machine learning will require information about the syntax of medical language. Understanding the syntactic differences between medical and non-medical corpora may allow more efficient acquisition of a lexicon. Three experiments designed to quantify the syntactic differences in medical and non-medical corpora were conducted. The results show that the syntax of medical language shows less variation than non-medical language and is likely simpler. The differences were great enough to question the applicability of general language tools on medical language. These differences may reduce the difficulty of some free text machine learning problems by capitalizing on the simpler nature of narrative medical syntax.

Artificial Intelligence↗

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗

Molecular hashkeys: a novel method for molecular characterization and its application for predicting important pharmaceutical properties of molecules.

We define a novel numerical molecular representation, called the molecular hashkey, that captures sufficient information about a molecule to predict pharmaceutically interesting properties directly from three-dimensional molecular structure. The molecular hashkey represents molecular surface properties as a linear array of pairwise surface-based comparisons of the target molecule against a common 'basis-set' of molecules. Hashkey-measured molecular similarity correlates well with direct methods of measuring molecular surface similarity. Using a simple machine-learning technique with the molecular hashkeys, we show that it is possible to accurately predict the octanol-water partition coefficient, log P. Using more sophisticated learning techniques, we show that an accurate model of intestinal absorption for a set of drugs can be constructed using the same hashkeys used in the aforementioned experiments. Once a set of molecular hashkeys is calculated, its use in the training and testing of property-based models is very fast. Further, the required amount of data for model construction is very small. Neural network-based hashkey models trained on data sets as small as 30 molecules yield statistically significant prediction of molecular properties. The lack of a requirement for large data sets lends itself well to the prediction of pharmaceutically relevant molecular parameters for which data generation is expensive and slow. Molecular hashkeys coupled with machine-learning techniques can yield models that predict key pharmacological aspects of biologically important molecules and should therefore be important in the design of effective therapeutics.

Drug Design↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy.

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Databases, Protein↗

An Exosomal Signature for Preoperative Detection of Occult Liver Metastasis in Pancreatic Cancer.

IMPORTANCE: Early liver metastasis (early-LiM) after pancreatectomy represents an aggressive biological phenotype of pancreatic ductal adenocarcinoma (PDAC) and is associated with markedly poor survival. Reliable preoperative biomarkers to identify occult hepatic micrometastasis remain lacking. OBJECTIVE: To develop and externally validate a circulating exosomal microRNA (exo-miRNA)-based machine learning model for preoperative detection of occult early-LiM in PDAC. DESIGN, SETTING, AND PARTICIPANTS: This multicenter retrospective case-control study included 3 phases: genome-wide discovery using exo-miRNA sequencing (discovery cohort), model development (training cohort), and independent external validation (2 validation cohorts). The study took place at 4 medical centers in China, Japan, and South Korea. A total of 372 patients were enrolled between 2011 and 2024. Data were analyzed from July 2024 to November 2025. EXPOSURES: Circulating plasma-derived exosomal miRNA expression profiles. MAIN OUTCOMES AND MEASURES: The primary outcome was early-LiM, defined as liver recurrence within 6 months after curative-intent resection. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and survival outcomes were assessed using Kaplan-Meier analysis. RESULTS: Among 372 patients with PDAC (median [IQR] age, 67 [59-73] years; 229 [61.6%] male and 143 [38.4%] female; median follow-up among survivors, 969 days),early-LiM was associated with significantly worse overall survival compared with other recurrence patterns (median OS, 9.1 months vs 26.6-31.8 months; log-rank P&#x2009;<&#x2009;.001). A 7-exo-miRNA extreme gradient boosting model demonstrated discrimination in the training cohort (AUC, 0.899; 95% CI, 0.822-0.976) and maintained performance in external testing cohorts (AUC, 0.876; 95% CI, 0.846-0.951 and AUC, 0.862; 95% CI, 0.744-0.981). The exo-miRNA panel score remained an independent identifier of early-LiM in multivariable analysis (odds ratio, 26.49; 95% CI, 18.45-55.28; P&#x2009;<&#x2009;.001) and stratified overall survival (log-rank P&#x2009;<&#x2009;.001). Decision curve analysis suggested improved net clinical benefit compared with conventional clinicopathologic variables. CONCLUSION AND RELEVANCE: In this multicenter study, a circulating exo-miRNA-based machine learning model enabled preoperative detection of occult early liver metastasis risk in PDAC. These findings support the potential of exosomal biomarkers to inform biology-guided treatment sequencing and warrant prospective validation.

Journal Article↗

Senescent fibroblasts drive CD8+ T cell dysfunction in colorectal cancer via CD36-mediated lipid transfer and peroxidation.

BACKGROUND: Functional exhaustion of tumor-infiltrating CD8+ T cells represents a hallmark of colorectal cancer (CRC) immunosuppression, though its mechanistic drivers remain elusive. Given the established correlation between CRC progression and stromal senescence characterized by pathological lipid accumulation and impaired immunity, we investigated whether and how senescent fibroblasts actively regulate CD8+ T cell dysfunction. METHODS: Single-cell RNA sequencing (scRNA-seq) analysis was conducted to unveil the diverse fibroblast populations and the significant lipid metabolism changes between senescent fibroblasts and non-senescent fibroblasts in human CRC specimens and adjacent normal mucosa. Machine-learning identified senescent fibroblasts with a distinct gene signature. Cell-cell communication analysis was used to evaluate the interactions between senescent fibroblasts and CD8+ T cells in colorectal cancer. Co-culture experiments were conducted among senescent fibroblasts, CD8+ T cells and patient-derived organoids of CRC (CRC-PDOs), with the results evaluated with high-content imaging and propidium iodide/Hoechst 33,342 staining. Flow cytometry, ELISA and lipid pulse-chase with BODIPY FL C16 were performed to detect the alterations of CD8+ T cell cytotoxic function and metabolic status. AOM/DSS-induced CRC mouse model was used to conduct in vivo validation to evaluate whether senolytics could suppress CRC progression. Patients from the Cancer Genome Atlas colorectal cancer cohort were stratified into CD36-high and CD36-low groups by median expression, and drug sensitivity for GDSC2 compounds was predicted computationally using the oncoPredict R package. RESULTS: ScRNA-seq demonstrated the specific cell population presence and divergence of senescent fibroblasts between neoplastic and histologically normal adjacent cell clusters in CRC. Random Forest was employed for cell senescence classification. Feature importance analysis identified five genes as key contributors to the model&#x2019;s decision process. Cell-cell communication analysis revealed enhanced interactions between senescent fibroblasts and CD8+ T cells in CRC. Co-culture of senescent fibroblasts significantly impaired the cytotoxic functions of CD8+ T cells on CRC-PDOs, which was reflected by the declined proportions of granzyme B (GZMB) + and interferon gamma (IFN&#x3b3;) + CD8+ T cells and enhanced viability of CRC-PDOs. Mechanistically, the co-culture with senescent fibroblasts promoted the lipid shuttling into CD8+ T cells to induce lipid peroxidation and downstream impairment of cytotoxicity. Furthermore, the inhibition of CD36, the specific scavenger receptor for lipid uptake of CD8+ T cells, effectively suppressed lipid transfer and peroxidation thereby preserving the effector functions of CD8+ T cells and ultimately promoting tumor apoptosis. Complementarily, in vivo senolytic treatment significantly suppressed CRC progression in AOM-DSS CRC mouse models. Top 12 therapeutic agents were identified significantly enhanced predicted efficacy in CD36-high tumors. CONCLUSIONS: Our study identified a substantial population of senescent fibroblasts in human CRC through single cell transcriptomics, machine-learning and clinical biopsies. These senescent fibroblasts impair CD8+ T cell-mediated killing of CRC-PDOs via CD36-dependent lipid transfer, suggesting senolytic targeting of stromal cells as a promising immunotherapeutic strategy for CRC.

Colorectal Neoplasms↗