Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A regulatory network underlying idiopathic pulmonary fibrosis.

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive interstitial lung disease in which genetic susceptibility interacts with epithelial, immune, and mesenchymal remodeling. Although the chromosome 11p15.5 locus contains established IPF susceptibility signals near MUC5B and TOLLIP, the broader regulatory architecture of this region remains incompletely resolved. METHODS: We integrated IPF genome-wide association study summary statistics with methylation, expression, and protein quantitative trait loci using summary-data-based Mendelian randomization (SMR). SMR-prioritized candidates were evaluated in independent transcriptomic and methylation cohorts and further contextualized using microRNA, transcription-factor, protein-interaction, machine-learning, single-cell, and spatial transcriptomic analyses. Fibrosis-associated expression patterns were assessed in a bleomycin-induced pulmonary fibrosis rat model. RESULTS: The analyses recovered the established MUC5B and TOLLIP signals and prioritized BRSK2 as a comparatively underexplored candidate supported by eQTL-based SMR and independent molecular evidence. The BRSK2 pQTL association did not pass the HEIDI test and was therefore not interpreted as convergent protein-level genetic evidence. Network analyses linked BRSK2 to cell-cycle, metabolic-stress, and senescence-related programs, while cross-cohort machine learning prioritized FOXA2, CDC25B, and NFE2 as informative network features. Single-cell and spatial analyses localized BRSK2 preferentially to fibroblast and myofibroblast compartments and to regions with greater histological fibrosis severity. In fibrotic rat lungs, BRSK2 expression increased, whereas FOXA2 and CDC25B decreased at the transcript and protein levels. CONCLUSIONS: These findings refine the molecular landscape of the chromosome 11p15.5 IPF susceptibility locus and prioritize BRSK2 as a candidate component of an IPF-associated profibrotic fibroblast state. Its causal contribution, direct regulatory relationships, and therapeutic tractability require targeted mechanistic validation.

Idiopathic Pulmonary Fibrosis↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Predictive models for protein crystallization.

Crystallization of proteins is a nontrivial task, and despite the substantial efforts in robotic automation, crystallization screening is still largely based on trial-and-error sampling of a limited subset of suitable reagents and experimental parameters. Funding of high throughput crystallography pilot projects through the NIH Protein Structure Initiative provides the opportunity to collect crystallization data in a comprehensive and statistically valid form. Data mining and machine learning algorithms thus have the potential to deliver predictive models for protein crystallization. However, the underlying complex physical reality of crystallization, combined with a generally ill-defined and sparsely populated sampling space, and inconsistent scoring and annotation make the development of predictive models non-trivial. We discuss the conceptual problems, and review strengths and limitations of current approaches towards crystallization prediction, emphasizing the importance of comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high throughput crystallography initiatives, exchange of information and standardization should be encouraged, aiming to effectively integrate data mining and machine learning efforts into a comprehensive predictive framework for protein crystallization. Similar experimental design and knowledge discovery strategies should be applied to valid analysis and prediction of protein expression, solubilization, and purification, as well as crystal handling and cryo-protection.

Bayes Theorem↗

Molecular hashkeys: a novel method for molecular characterization and its application for predicting important pharmaceutical properties of molecules.

We define a novel numerical molecular representation, called the molecular hashkey, that captures sufficient information about a molecule to predict pharmaceutically interesting properties directly from three-dimensional molecular structure. The molecular hashkey represents molecular surface properties as a linear array of pairwise surface-based comparisons of the target molecule against a common 'basis-set' of molecules. Hashkey-measured molecular similarity correlates well with direct methods of measuring molecular surface similarity. Using a simple machine-learning technique with the molecular hashkeys, we show that it is possible to accurately predict the octanol-water partition coefficient, log P. Using more sophisticated learning techniques, we show that an accurate model of intestinal absorption for a set of drugs can be constructed using the same hashkeys used in the aforementioned experiments. Once a set of molecular hashkeys is calculated, its use in the training and testing of property-based models is very fast. Further, the required amount of data for model construction is very small. Neural network-based hashkey models trained on data sets as small as 30 molecules yield statistically significant prediction of molecular properties. The lack of a requirement for large data sets lends itself well to the prediction of pharmaceutically relevant molecular parameters for which data generation is expensive and slow. Molecular hashkeys coupled with machine-learning techniques can yield models that predict key pharmacological aspects of biologically important molecules and should therefore be important in the design of effective therapeutics.

Drug Design↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy.

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Databases, Protein↗

An Exosomal Signature for Preoperative Detection of Occult Liver Metastasis in Pancreatic Cancer.

IMPORTANCE: Early liver metastasis (early-LiM) after pancreatectomy represents an aggressive biological phenotype of pancreatic ductal adenocarcinoma (PDAC) and is associated with markedly poor survival. Reliable preoperative biomarkers to identify occult hepatic micrometastasis remain lacking. OBJECTIVE: To develop and externally validate a circulating exosomal microRNA (exo-miRNA)-based machine learning model for preoperative detection of occult early-LiM in PDAC. DESIGN, SETTING, AND PARTICIPANTS: This multicenter retrospective case-control study included 3 phases: genome-wide discovery using exo-miRNA sequencing (discovery cohort), model development (training cohort), and independent external validation (2 validation cohorts). The study took place at 4 medical centers in China, Japan, and South Korea. A total of 372 patients were enrolled between 2011 and 2024. Data were analyzed from July 2024 to November 2025. EXPOSURES: Circulating plasma-derived exosomal miRNA expression profiles. MAIN OUTCOMES AND MEASURES: The primary outcome was early-LiM, defined as liver recurrence within 6 months after curative-intent resection. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and survival outcomes were assessed using Kaplan-Meier analysis. RESULTS: Among 372 patients with PDAC (median [IQR] age, 67 [59-73] years; 229 [61.6%] male and 143 [38.4%] female; median follow-up among survivors, 969 days),early-LiM was associated with significantly worse overall survival compared with other recurrence patterns (median OS, 9.1 months vs 26.6-31.8 months; log-rank P&#x2009;<&#x2009;.001). A 7-exo-miRNA extreme gradient boosting model demonstrated discrimination in the training cohort (AUC, 0.899; 95% CI, 0.822-0.976) and maintained performance in external testing cohorts (AUC, 0.876; 95% CI, 0.846-0.951 and AUC, 0.862; 95% CI, 0.744-0.981). The exo-miRNA panel score remained an independent identifier of early-LiM in multivariable analysis (odds ratio, 26.49; 95% CI, 18.45-55.28; P&#x2009;<&#x2009;.001) and stratified overall survival (log-rank P&#x2009;<&#x2009;.001). Decision curve analysis suggested improved net clinical benefit compared with conventional clinicopathologic variables. CONCLUSION AND RELEVANCE: In this multicenter study, a circulating exo-miRNA-based machine learning model enabled preoperative detection of occult early liver metastasis risk in PDAC. These findings support the potential of exosomal biomarkers to inform biology-guided treatment sequencing and warrant prospective validation.

Journal Article↗

Senescent fibroblasts drive CD8+ T cell dysfunction in colorectal cancer via CD36-mediated lipid transfer and peroxidation.

BACKGROUND: Functional exhaustion of tumor-infiltrating CD8+ T cells represents a hallmark of colorectal cancer (CRC) immunosuppression, though its mechanistic drivers remain elusive. Given the established correlation between CRC progression and stromal senescence characterized by pathological lipid accumulation and impaired immunity, we investigated whether and how senescent fibroblasts actively regulate CD8+ T cell dysfunction. METHODS: Single-cell RNA sequencing (scRNA-seq) analysis was conducted to unveil the diverse fibroblast populations and the significant lipid metabolism changes between senescent fibroblasts and non-senescent fibroblasts in human CRC specimens and adjacent normal mucosa. Machine-learning identified senescent fibroblasts with a distinct gene signature. Cell-cell communication analysis was used to evaluate the interactions between senescent fibroblasts and CD8+ T cells in colorectal cancer. Co-culture experiments were conducted among senescent fibroblasts, CD8+ T cells and patient-derived organoids of CRC (CRC-PDOs), with the results evaluated with high-content imaging and propidium iodide/Hoechst 33,342 staining. Flow cytometry, ELISA and lipid pulse-chase with BODIPY FL C16 were performed to detect the alterations of CD8+ T cell cytotoxic function and metabolic status. AOM/DSS-induced CRC mouse model was used to conduct in vivo validation to evaluate whether senolytics could suppress CRC progression. Patients from the Cancer Genome Atlas colorectal cancer cohort were stratified into CD36-high and CD36-low groups by median expression, and drug sensitivity for GDSC2 compounds was predicted computationally using the oncoPredict R package. RESULTS: ScRNA-seq demonstrated the specific cell population presence and divergence of senescent fibroblasts between neoplastic and histologically normal adjacent cell clusters in CRC. Random Forest was employed for cell senescence classification. Feature importance analysis identified five genes as key contributors to the model&#x2019;s decision process. Cell-cell communication analysis revealed enhanced interactions between senescent fibroblasts and CD8+ T cells in CRC. Co-culture of senescent fibroblasts significantly impaired the cytotoxic functions of CD8+ T cells on CRC-PDOs, which was reflected by the declined proportions of granzyme B (GZMB) + and interferon gamma (IFN&#x3b3;) + CD8+ T cells and enhanced viability of CRC-PDOs. Mechanistically, the co-culture with senescent fibroblasts promoted the lipid shuttling into CD8+ T cells to induce lipid peroxidation and downstream impairment of cytotoxicity. Furthermore, the inhibition of CD36, the specific scavenger receptor for lipid uptake of CD8+ T cells, effectively suppressed lipid transfer and peroxidation thereby preserving the effector functions of CD8+ T cells and ultimately promoting tumor apoptosis. Complementarily, in vivo senolytic treatment significantly suppressed CRC progression in AOM-DSS CRC mouse models. Top 12 therapeutic agents were identified significantly enhanced predicted efficacy in CD36-high tumors. CONCLUSIONS: Our study identified a substantial population of senescent fibroblasts in human CRC through single cell transcriptomics, machine-learning and clinical biopsies. These senescent fibroblasts impair CD8+ T cell-mediated killing of CRC-PDOs via CD36-dependent lipid transfer, suggesting senolytic targeting of stromal cells as a promising immunotherapeutic strategy for CRC.

Colorectal Neoplasms↗

Prediction of P-glycoprotein substrates by a support vector machine approach.

P-glycoproteins (P-gp) actively transport a wide variety of chemicals out of cells and function as drug efflux pumps that mediate multidrug resistance and limit the efficacy of many drugs. Methods for facilitating early elimination of potential P-gp substrates are useful for facilitating new drug discovery. A computational ensemble pharmacophore model has recently been used for the prediction of P-gp substrates with a promising accuracy of 63%. It is desirable to extend the prediction range beyond compounds covered by the known pharmacophore models. For such a purpose, a machine learning method, support vector machine (SVM), was explored for the prediction of P-gp substrates. A set of 201 chemical compounds, including 116 substrates and 85 nonsubstrates of P-gp, was used to train and test a SVM classification system. This SVM system gave a prediction accuracy of at least 81.2% for P-gp substrates based on two different evaluation methods, which is substantially improved against that obtained from the multiple-pharmacophore model. The prediction accuracy for nonsubstrates of P-gp is 79.2% using 5-fold cross-validation. These accuracies are slightly better than those obtained from other statistical classification methods, including k-nearest neighbor (k-NN), probabilistic neural networks (PNN), and C4.5 decision tree, that use the same sets of data and molecular descriptors. Our study indicates the potential of SVM in facilitating the prediction of P-gp substrates.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

BCI Competition 2003--Data set IIb: support vector machines for the P300 speller paradigm.

We propose an approach to analyze data from the P300 speller paradigm using the machine-learning technique support vector machines. In a conservative classification scheme, we found the correct solution after five repetitions. While the classification within the competition is designed for offline analysis, our approach is also well-suited for a real-world online solution: It is fast, requires only 10 electrode positions and demands only a small amount of preprocessing.

Algorithms↗

NR3C1 Modulates Wnt Signalling to Influence the Invasiveness and Immune Features of Nonfunctioning Invasive Pituitary Adenomas.

Pituitary adenomas (PAs) are common intracranial tumours, and invasiveness in nonfunctioning invasive pituitary adenomas (NIPAs) predicts poor prognosis. The molecular mechanisms driving this phenotype remain unclear. This study explored the role of nuclear receptor subfamily 3 group C member 1 (NR3C1) in NIPA invasiveness and its regulation of Wnt signalling. mRNA expression profiles of 32 PA samples were generated by RNA-seq, and proteomic data from 19 samples were obtained by mass spectrometry. Immune-related differentially expressed genes (DEGs) were retrieved from GeneCards. Weighted gene coexpression network analysis identified modules and hub genes linked to invasiveness, while machine learning methods (support vector machine, LASSO, random forest) prioritised key genes. Gene set enrichment analysis (GSEA) assessed pathways associated with candidate gene expression. NR3C1 expression and function were validated by immunohistochemistry, Western blotting and invasion assays. Integration of transcriptomic, proteomic and immune-related datasets yielded 11 overlapping genes, with NR3C1 emerging as the top candidate. NR3C1 was significantly upregulated in NIPAs and demonstrated good discriminatory power by ROC analysis. GSEA associated high NR3C1 expression with Wnt pathway activation. Functional experiments confirmed that NR3C1 overexpression enhances the invasive capacity of PA cells. NR3C1 promotes the invasive phenotype of NIPAs by activating Wnt signalling. These findings suggest NR3C1 as a potential biomarker and therapeutic target for invasive pituitary adenomas.

Humans↗

Prediction of RNA-binding proteins from primary sequence by a support vector machine approach.

Elucidation of the interaction of proteins with different molecules is of significance in the understanding of cellular processes. Computational methods have been developed for the prediction of protein-protein interactions. But insufficient attention has been paid to the prediction of protein-RNA interactions, which play central roles in regulating gene expression and certain RNA-mediated enzymatic processes. This work explored the use of a machine learning method, support vector machines (SVM), for the prediction of RNA-binding proteins directly from their primary sequence. Based on the knowledge of known RNA-binding and non-RNA-binding proteins, an SVM system was trained to recognize RNA-binding proteins. A total of 4011 RNA-binding and 9781 non-RNA-binding proteins was used to train and test the SVM classification system, and an independent set of 447 RNA-binding and 4881 non-RNA-binding proteins was used to evaluate the classification accuracy. Testing results using this independent evaluation set show a prediction accuracy of 94.1%, 79.3%, and 94.1% for rRNA-, mRNA-, and tRNA-binding proteins, and 98.7%, 96.5%, and 99.9% for non-rRNA-, non-mRNA-, and non-tRNA-binding proteins, respectively. The SVM classification system was further tested on a small class of snRNA-binding proteins with only 60 available sequences. The prediction accuracy is 40.0% and 99.9% for snRNA-binding and non-snRNA-binding proteins, indicating a need for a sufficient number of proteins to train SVM. The SVM classification systems trained in this work were added to our Web-based protein functional classification software SVMProt, at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi. Our study suggests the potential of SVM as a useful tool for facilitating the prediction of protein-RNA interactions.

Algorithms↗

Honest assessments of automatic learning algorithm performance.

OBJECTIVE: To compare methods of evaluating probabilistic predictors in systems that learn from examples. STUDY DESIGN: The performance of four automatic learning algorithms, representing current machine learning technology, were assessed using four methodologies in the task of separating normal squamous intermediate cervical cells from all other segmented objects in digital images. Two of the methodologies were carefully constructed to model sources of variation associated with the choice of training and test sets. These assessments were statistically compared with assessments using both standard and a modified version of cross-validation. RESULTS: The investigation illustrates the tradeoffs involved in obtaining statistical rigor as compared with the cost of collecting data. While cross-validation makes frugal use of data, it can produce misleading assessments of algorithm performance in terms of both bias and variance. The modified version produces more reliable assessments but in some cases may also be misleading. CONCLUSION: We suggest that users of learning algorithms should exercise judicious care in evaluating learning algorithm performance in order to avoid unnecessary bias and large variance in their assessments.

Algorithms↗