Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Feature subset selection by genetic algorithms and estimation of distribution algorithms. A case study in the survival of cirrhotic patients treated with TIPS.

The transjugular intrahepatic portosystemic shunt (TIPS) is an interventional treatment for cirrhotic patients with portal hypertension. In the light of our medical staff's experience, the consequences of TIPS are not homogeneous for all the patients and a subgroup dies in the first 6 months after TIPS placement. Actually, there is no risk indicator to identify this subgroup of patients before treatment. An investigation for predicting the survival of cirrhotic patients treated with TIPS is carried out using a clinical database with 107 cases and 77 attributes. Four supervised machine learning classifiers are applied to discriminate between both subgroups of patients. The application of several feature subset selection (FSS) techniques has significantly improved the predictive accuracy of these classifiers and considerably reduced the amount of attributes in the classification models. Among FSS techniques, FSS-TREE, a new randomized algorithm inspired on the new EDA (estimation of distribution algorithm) paradigm has obtained the best average accuracy results for each classifier.

Algorithms↗

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals↗

ASGCL: Adaptive Sparse Mapping-based graph contrastive learning network for cancer drug response prediction.

Personalized cancer drug treatment is emerging as a frontier issue in modern medical research. Considering the genomic differences among cancer patients, determining the most effective drug treatment plan is a complex and crucial task. In response to these challenges, this study introduces the Adaptive Sparse Graph Contrastive Learning Network (ASGCL), an innovative approach to unraveling latent interactions in the complex context of cancer cell lines and drugs. The core of ASGCL is the GraphMorpher module, an innovative component that enhances the input graph structure via strategic node attribute masking and topological pruning. By contrasting the augmented graph with the original input, the model delineates distinct positive and negative sample sets at both node and graph levels. This dual-level contrastive approach significantly amplifies the model's discriminatory prowess in identifying nuanced drug responses. Leveraging a synergistic combination of supervised and contrastive loss, ASGCL accomplishes end-to-end learning of feature representations, substantially outperforming existing methodologies. Comprehensive ablation studies underscore the efficacy of each component, corroborating the model's robustness. Experimental evaluations further illuminate ASGCL's proficiency in predicting drug responses, offering a potent tool for guiding clinical decision-making in cancer therapy.

Humans↗

Bio-inspired computing tissues: towards machines that evolve, grow, and learn.

Biological inspiration in the design of computing machines could allow the creation of new machines with promising characteristics such as fault-tolerance, self-replication or cloning, reproduction, evolution, adaptation and learning, and growth. The aim of this paper is to introduce bio-inspired computing tissues that might constitute a key concept for the implementation of 'living' machines. We first present a general overview of bio-inspired systems and the POE model that classifies bio-inspired machines along three axes. The Embryonics project--inspired by some of the basic processes of molecular biology--is described by means of the BioWatch application, a fault-tolerant and self-repairable watch. The main characteristics of the Embryonics project are the multicellular organization, the cellular differentiation, and the self-repair capabilities. The BioWall is intended as a reconfigurable computing tissue, capable of interacting with its environment by means of a large number of touch-sensitive elements coupled with a color display. For illustrative purposes, a large-scale implementation of the BioWatch on the BioWall's computational tissue is presented. We conclude the paper with a description of bio-inspired computing tissues and POEtic machines.

Cell Differentiation↗

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified ∼218,000 OGs supported by expression evidence, while ∼330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

Causal associations between hormone replacement therapy and brain structure: Evidence from large-scale Mendelian randomization and double machine learning.

BACKGROUND: Hormone replacement therapy (HRT) is widely prescribed for the management of hormone deficiency, particularly during menopause, yet its causal effects on human brain structure remain incompletely understood. Observational studies have reported heterogeneous associations, underscoring the need for robust causal inference. METHODS: We applied an integrated causal framework combining two-sample Mendelian Randomization (MR) and Double Machine Learning (DML) to evaluate the effects of four HRT-related exposures-age at initiation, age at cessation, ever-use of HRT, and a composite medication-based phenotype-on 1366 brain imaging-derived phenotypes from the UK Biobank. Genetic instruments were derived from large-scale GWAS summary statistics, and causal estimates were validated using non-parametric DML models with cross-fitting and performance evaluation. RESULTS: Genetic instruments for age at HRT initiation, age at cessation, and ever-use of HRT were strong (median F-statistics 16.29-36.66). MR analyses identified a causal association between later initiation of HRT and lower orientation dispersion in the right inferior cerebellar peduncle (ubm-a-542; primary finding, no pleiotropy detected). An additional association with the left tapetum FA (ubm-a-243) was identified but exhibited significant directional horizontal pleiotropy (MR-Egger intercept P = 0.001) and is excluded from primary conclusions (Supplementary Note S2). Later cessation of HRT was associated with increased cortical thickness in the left middle occipital gyrus, reduced surface area in the left frontopolar cortex, and increased orientation dispersion in the splenium of the corpus callosum. Ever-use of HRT was causally linked to larger volumes of the right inferior frontal gyrus and right nucleus accumbens. These associations were corroborated by independent DML validation, which provided causally debiased estimates robust to high-dimensional confounding. Results for ukb-b-8080 (median F = 1.45) are provided in Supplementary Note S1 only; weak-instrument bias precludes causal inference. CONCLUSIONS: This study provides genetic-instrument-based and machine-learning-validated evidence for causal associations between HRT exposure-particularly its timing and lifetime use-and specific features of human brain structure, including white-matter microarchitecture, cortical thickness, and regional brain volume. These findings are FDR-controlled within exposures and independently replicated by DML, but require replication in external neuroimaging GWAS cohorts to establish definitive causal conclusions. They highlight the neurobiological relevance of sex steroid exposure and inform future research on brain aging and personalized hormone-based interventions.

Humans↗

A data-mining approach to spacer oligonucleotide typing of Mycobacterium tuberculosis.

MOTIVATION: The Direct Repeat (DR) locus of Mycobacterium tuberculosis is a suitable model to study (i) molecular epidemiology and (ii) the evolutionary genetics of tuberculosis. This is achieved by a DNA analysis technique (genotyping), called sp acer oligo nucleotide typing (spoligotyping ). In this paper, we investigated data analysis methods to discover intelligible knowledge rules from spoligotyping, that has not yet been applied on such representation. This processing was achieved by applying the C4.5 induction algorithm and knowledge rules were produced. Finally, a Prototype Selection (PS) procedure was applied to eliminate noisy data. This both simplified decision rules, as well as the number of spacers to be tested to solve classification tasks. In the second part of this paper, the contribution of 25 new additional spacers and the knowledge rules inferred were studied from a machine learning point of view. From a statistical point of view, the correlations between spacers were analyzed and suggested that both negative and positive ones may be related to potential structural constraints within the DR locus that may shape its evolution directly or indirectly. RESULTS: By generating knowledge rules induced from decision trees, it was shown that not only the expert knowledge may be modeled but also improved and simplified to solve automatic classification tasks on unknown patterns. A practical consequence of this study may be a simplification of the spoligotyping technique, resulting in a reduction of the experimental constraints and an increase in the number of samples processed.

Algorithms↗

EDAmame: interactive exploratory data analyses with explainable models.

SUMMARY: Complex tabular datasets comprising many diverse features can require specific expertise to interpret, posing a barrier to researchers with minimal data science experience. EDAmame is an interactive tool that simplifies initial analysis and visualization of these datasets, providing insights into data quality and feature relationships. By leveraging open-source machine learning frameworks in R, EDAmame allows researchers to perform effective exploratory data analysis without command-line or coding requirements. AVAILABILITY AND IMPLEMENTATION: A limited online version can be accessed at https://edamame.org.au/ or can be downloaded from https://doi.org/10.5281/zenodo.15356492. The app is developed in R Shiny and implements tidyverse and tidymodels packages.

Machine Learning↗

Improving the efficiency of a user-driven learning system with reconfigurable hardware. Application to DNA splicing.

This paper describes a new approach to problem solving by splitting up problem component parts between software and hardware. Our main idea arises from the combination of two previously published works. The first one proposed a conceptual environment of concept modelling in which the machine and the human expert interact. The second one reported an algorithm based on reconfigurable hardware system which outperforms any kind of previously published genetic data base scanning hardware or algorithms. Here we show how efficient the interaction between the machine and the expert is when the concept modelling is based on reconfigurable hardware system. Their cooperation is thus achieved with an real time interaction speed. The designed system has been partially applied to the recognition of primate splice junctions sites in genetic sequences.

Algorithms↗

Statistical mechanics of learning with soft margin classifiers.

We study the typical learning properties of the recently introduced soft margin classifiers (SMCs), learning realizable and unrealizable tasks, with the tools of statistical mechanics. We derive analytically the behavior of the learning curves in the regime of very large training sets. We obtain exponential and power laws for the decay of the generalization error towards the asymptotic value, depending on the task and on general characteristics of the distribution of stabilities of the patterns to be learned. The optimal learning curves of the SMCs, which give the minimal generalization error, are obtained by tuning the coefficient controlling the trade-off between the error and the regularization terms in the cost function. If the task is realizable by the SMC, the optimal performance is better than that of a hard margin support vector machine and is very close to that of a Bayesian classifier.

Algorithms↗

Machine learning in quantitative histopathology.

The role of expert systems functioning as process controllers in learning image understanding systems is discussed. Numeric learning systems already have found a number of applications in cytologic and histopathologic diagnosis. Depending on the required capabilities, systems of increasing complexity are needed. Expert systems to guide scene segmentation in histopathologic imagery require model-based reasoning. Diagnostic image interpretation with learning capability demands a full model of the human expert's competence, including a considerable variety of knowledge representation schemes and inference strategies, coordinated by a meta-process controller.

Artificial Intelligence↗

Multi-cohort integration and machine learning identify CPVL as a novel oncogenic driver in gastric cancer.

BACKGROUND: Gastric cancer (GC) remains a leading cause of cancer-related mortality worldwide, and the prognosis of advanced GC remains poor. Systematic identification of robust biomarkers through multi-cohort integration and computational prioritization may facilitate the discovery of novel therapeutic targets. AIM: To identify key genes associated with gastric cancer progression through integrative multi-omics analysis and to elucidate the biological functions and molecular mechanisms of the top-prioritized candidate gene. METHODS: Comprehensive bioinformatics analyses integrating The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Gene Expression Omnibus (GEO) datasets were performed using differential expression analysis, weighted gene co-expression network analysis (WGCNA), Cox regression, and eight machine-learning algorithms to systematically identify and prioritize GC-associated hub genes. Among the identified candidates, CPVL was selected for further validation based on its diagnostic and prognostic performance. CPVL expression and clinical relevance were validated by independent datasets and immunohistochemistry. Lentiviral constructs were used to overexpress or silence CPVL in GC cell lines. Functional assays were performed, including CCK-8, colony formation, EdU incorporation, and flow cytometry, to assess cell proliferation and cell-cycle distribution. Western blotting and JAK2 inhibitor (AZD1480) rescue experiments were performed to elucidate the underlying mechanisms, and a nude mouse xenograft model was used to evaluate tumorigenicity in vivo. RESULTS: Multi-cohort screening identified five hub genes (CPVL, AADAC, BCAT1, CPXM1, and FBN1). Among them, CPVL exhibited the highest diagnostic accuracy (AUC = 0.895) and the strongest correlation with poor overall survival, and was therefore selected for mechanistic investigation. CPVL expression was markedly upregulated in GC tissues and cell lines. Functional assays demonstrated that CPVL promotes GC cell proliferation and accelerates G1/S-phase transition. Mechanistically, CPVL activated the JAK2/STAT3 signaling pathway, upregulating Cyclin D1 and CDK4 while downregulating p27. Treatment with the JAK2 inhibitor AZD1480 partially reversed these effects. In vivo, CPVL knockdown significantly inhibited tumor growth. CONCLUSION: Through systematic multi-cohort integration and machine-learning prioritization, CPVL was identified as a novel oncogenic driver in gastric cancer. CPVL promotes tumor growth via activation of the JAK2/STAT3 pathway and regulation of the Cyclin D1/CDK4/p27 axis, highlighting its potential as a diagnostic biomarker and therapeutic target.

Biomarker↗

A time-resolved single-cell roadmap of the logic driving anterior neural crest diversification from neural border to migration stages.

Neural crest cells exemplify cellular diversification from a multipotent progenitor population. However, the full sequence of early molecular choices orchestrating the emergence of neural crest heterogeneity from the embryonic ectoderm remains elusive. Gene-regulatory-networks (GRN) govern early development and cell specification toward definitive neural crest. Here, we combine ultradense single-cell transcriptomes with machine-learning and large-scale transcriptomic and epigenomic experimental validation of selected trajectories, to provide the general principles and highlight specific features of the GRN underlying neural crest fate diversification from induction to early migration stages using Xenopus frog embryos as a model. During gastrulation, a transient neural border zone state precedes the choice between neural crest and placodes which includes multiple converging gene programs. During neurulation, transcription factor connectome, and bifurcation analyses demonstrate the early emergence of neural crest fates at the neural plate stage, alongside an unbiased multipotent-like lineage persisting until epithelial-mesenchymal transition stage. We also decipher circuits driving cranial and vagal neural crest formation and provide a broadly applicable high-throughput validation strategy for investigating single-cell transcriptomes in vertebrate GRNs in development, evolution, and disease.

Animals↗

The automatic discovery of structural principles describing protein fold space.

The study of protein structure has been driven largely by the careful inspection of experimental data by human experts. However, the rapid determination of protein structures from structural-genomics projects will make it increasingly difficult to analyse (and determine the principles responsible for) the distribution of proteins in fold space by inspection alone. Here, we demonstrate a machine-learning strategy that automatically determines the structural principles describing 45 folds. The rules learnt were shown to be both statistically significant and meaningful to protein experts. With the increasing emphasis on high-throughput experimental initiatives, machine-learning and other automated methods of analysis will become increasingly important for many biological problems.

Algorithms↗

A case-acquisition and decision-support system for the analysis of group-average lactation curves.

A case-acquisition and decision-support system was developed to support the analysis of group-average lactation curves and to acquire example cases from domain specialists. This software was developed through several iterations of a three-step approach involving 1) problem analysis and formulation in consultation with two dairy nutrition specialists; 2) development of a case-acquisition and decision-support prototype by the system developer; and 3) use of the prototype by the domain specialists to analyze and classify milk-recording data from example herds. The overall problem was decomposed into three subproblems: removal of outlier tests and lactation curves of individual cows; interpretation of group-average lactation curves; and diagnosis of detected abnormalities at the herd level through the identification of potential management deficiencies. For each subproblem, a software module was developed allowing the user to analyze both graphical and numerical performance representations and classify these representations using predefined linguistic descriptors. The example-based method for the development of the program proved to be very useful, facilitating the communication between system developer and domain specialists, and allowing the specialists to explore the appropriateness of the various prototypes developed. The resulting software represents a formalization of the approach to group-average lactation curve analysis, elicited from the two domain specialists. In future research, the case-acquisition and decision-support system will be complemented with knowledge to automate identified classification tasks, which will be captured through the application of machine-learning techniques to example cases, acquired from domain specialists using the software.

Animals↗

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery↗

Discovering hidden candidate plastic-degrading enzymes: Combined multi-omics and machine learning strategy.

Plastic pollution poses a major threat to the stability of natural ecosystems as well as human health. Microbial enzymes have long been considered a potential resource for targeted biodegradation but, except for a few successful cases, the discovery of efficient enzymes has proved challenging. Aiming to accelerate the process, we propose an approach combining metagenomics, metatranscriptomics and semi-supervised learning that selects promising plastic-degrading candidate enzymes from the proteome of relevant microorganisms. Tested on a dataset of over 10,000 microbial proteins, ranking models consistently prioritize known plastic-degrading enzymes, achieving an area under the cumulative distribution function curve above 0.96, with leave-one-family-out cross-validation indicating that performance is largely retained across protein families. As a case study, this work focuses on mixed microbial cultures exposed for extended periods to polyethylene, polyethylene terephthalate, and polyurethane substrates. The prevalent species after selective enrichment were functionally characterized, finding Rhodococcus aetherivorans as the most relevant species in two of the five cultures under investigation. Among the top-ranked proteins, several have high structural similarity with known enzymes despite not being identified by sequence similarity search. Moreover, according to metatranscriptomics results, several of these enzymes were found to be expressed at the same level or above that of annotated enzymes, suggesting that they may have functional relevance. Overall, this work highlights the potential of integrating multi-omics with data-driven methods for enzyme discovery and for accelerating the development of biotechnological solutions to plastic pollution.

Biodegradation, Environmental↗