Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Consciousness in a Self-Learning, Memory-Controlled, Compound Machine.

A memory-controlled, sensor/actuator machine senses conditions in its environment at given moments, and attempts to produce an action based upon its memory. However, a sensor/actuator machine will stop producing new behavior if its environment is removed. A sensor/sensor unit can be added to the sensor/actuator machine, forming a compound machine. The sensor/sensor unit produces a stream of internally created sensed conditions, which can replace the sensed conditions from the environment. This illusion of an environment is similar to consciousness. In addition, actuator/sensor and actuator/actuator units can be added to this compound machine to further enhance its ability to function without an environment. Predetermined and empirical memory cells can be distributed throughout the control units of this compound machine to provide instinctive and learned behavior. The internal and exterior behavior of this compound machine can be modified greatly by changing the cycle start and ramp signals that activate these different kinds of memory cells. These signals are similar in form to brain waves.

Journal Article↗

Livestock Multi-Omics Integration: A Systematic Framework From Statistical Association to Causal Interpretation.

Livestock multi-omics integration is key to unraveling complex trait regulation, yet systematic, livestock-specific strategies remain scarce. This review traces the progression from single-omics accumulation to multi-dimensional integration, highlighting how large-scale genomic, epigenomic, and transcriptomic projects lay the foundation for functional dissection. We identify core impediments: extreme species diversity, marked data heterogeneity, limited sample sizes, and a pervasive reduction of multi-omics data to simplistic differential screens, resulting in low translational efficiency. We critically appraise four common pitfalls-overinterpreting correlation as causation, relegating proteomics to corroborating transcriptomics, incomplete microbiome-host integration lacking environmental context, and systematic neglect of metabolic fluxomics-and show how exposomics and fluxomics add necessary causal and dynamic dimensions. To address these, we propose a livestock-adapted three-tier analytical framework: (1) statistical association of cross-omics covariation patterns; (2) machine learning-driven feature mining and integrative modeling; and (3) causal interpretation encompassing Mendelian randomization, prior-knowledge-guided network inference, and physical causal evidence via fluxomics and metabolic control analysis. We further discuss how multimodal sequencing (single-cell, spatial, temporal) and generative AI can fundamentally mitigate heterogeneity and strengthen causal evidence. Finally, we outline future priorities in database standardization, livestock-specific benchmarking, and translational pipelines, charting a path from correlation-centric reporting to mechanistic causality and precision breeding.

Animals↗

Bio-inspired computing tissues: towards machines that evolve, grow, and learn.

Biological inspiration in the design of computing machines could allow the creation of new machines with promising characteristics such as fault-tolerance, self-replication or cloning, reproduction, evolution, adaptation and learning, and growth. The aim of this paper is to introduce bio-inspired computing tissues that might constitute a key concept for the implementation of 'living' machines. We first present a general overview of bio-inspired systems and the POE model that classifies bio-inspired machines along three axes. The Embryonics project--inspired by some of the basic processes of molecular biology--is described by means of the BioWatch application, a fault-tolerant and self-repairable watch. The main characteristics of the Embryonics project are the multicellular organization, the cellular differentiation, and the self-repair capabilities. The BioWall is intended as a reconfigurable computing tissue, capable of interacting with its environment by means of a large number of touch-sensitive elements coupled with a color display. For illustrative purposes, a large-scale implementation of the BioWatch on the BioWall's computational tissue is presented. We conclude the paper with a description of bio-inspired computing tissues and POEtic machines.

Cell Differentiation↗

Unveiling m7G modification patterns and causal drivers governing intracranial aneurysm rupture risk through multi-omics validation and m7G-MeRIP-seq profiling.

Intracranial aneurysm (IA) rupture causes severe brain hemorrhage with high mortality, yet its molecular drivers remain unclear and better risk prediction is urgently needed. Using transcriptomics, single-cell analysis, and genetic data, we investigated the role of N7-methylguanosine (m7G) RNA modification in IA. We identified distinct m7G modification patterns, validated their methylation features in patient samples, and incorporated these patterns into a machine learning-based rupture prediction model. The presence and characteristics of m7G patterns significantly improved model performance, achieving high predictive accuracy across three independent cohorts (AUC 0.91-0.95). Genetic analyses further identified three causal m7G-related genes (NSUN2, IFIT5, SNUPN), and laboratory experiments confirmed their altered expression and methylation in ruptured aneurysms. Overall, our findings demonstrate that m7G modifications play a key role in IA rupture. The validated prediction model offers strong clinical potential for rupture risk assessment, and the identified genes represent promising therapeutic targets.

Humans↗

Learning in higher order Boltzmann machines using linear response.

We introduce an efficient method for learning and inference in higher order Boltzmann machines. The method is based on mean field theory with the linear response correction. We compute the correlations using the exact and the approximated method for a fully connected third order network of ten neurons. In addition, we compare the results of the exact and approximate learning algorithm. Finally we use the presented method to solve the shifter problem. We conclude that the linear response approximation gives good results as long as the couplings are not too large.

Artificial Intelligence↗

Stable behavior in a recurrent neural network for a finite state machine.

For the learning of a finite state machine (FSM) by a recurrent neural network (RNN), we think about how to train an RNN so as to stably mimic an FSM even for sequences having a long length. First, we consider the relationship between the stable behavior and the internal representation of states, that is, clusters of the internal units' outputs. As for this relationship, we prove that an RNN can get the stable cluster transitions when a neuron activation parameter is larger than a certain finite value micro0. Secondly, to acquire the stable behavior, we regard the internal representation for the stable behavior as prior knowledge. This produces a new target function of learning with internal representation term. We derive a Bayesian style method to estimate coefficients of the terms in the function, corresponding to hyperparameters. Finally, experiments show that RNNs readily acquire stable behavior by using our proposed method.

Artificial Intelligence↗

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense↗

Universal learning curves of support vector machines.

Using methods of statistical physics, we investigate the role of model complexity in learning with support vector machines (SVMs), which are an important alternative to neural networks. We show the advantages of using SVMs with kernels of infinite complexity on noisy target rules, which, in contrast to common theoretical beliefs, are found to achieve optimal generalization error although the training error does not converge to the generalization error. Moreover, we find a universal asymptotics of the learning curves which depend only on the target rule but not on the SVM kernel.

Journal Article↗

Phylogenetic Methods Meet Deep Learning.

Deep learning (DL) has been widely used in various scientific fields, but its integration into phylogenetics has been slower, primarily due to the complex nature of phylogenetic data. The studies that apply DL to sequencing data often limit analyses to four-taxon trees. Many of these studies serve as "proof of principle" and perform similarly to traditional phylogeny reconstruction methods. New ways of using training data, such as encoding with compact bijective ladderized vectors or transformers, enable the handling of much larger trees and genomic data sets. This short perspective focuses on the application of DL in phylogenetics, introducing prevalent DL architectures. We highlight potential problems in the field by discussing the risks of using simulation-based training data and emphasize the importance of reproducibility and robustness in computational estimates. Finally, we explore promising research areas, including the combination of phylogenetics and population genetics in DL, the analysis of neighbor dependencies, and the potential to significantly reduce computational cost compared to traditional methods. This perspective illustrates the potential of DL in complementing traditional phylogeny reconstruction methods and aiding the advancement of phylogenetic analysis, especially in performing computationally demanding tasks such as model selection or estimating branch support values.

Humans↗

Kernel hierarchical gene clustering from microarray expression data.

MOTIVATION: Unsupervised analysis of microarray gene expression data attempts to find biologically significant patterns within a given collection of expression measurements. For example, hierarchical clustering can be applied to expression profiles of genes across multiple experiments, identifying groups of genes that share similar expression profiles. Previous work using the support vector machine supervised learning algorithm with microarray data suggests that higher-order features, such as pairwise and tertiary correlations across multiple experiments, may provide significant benefit in learning to recognize classes of co-expressed genes. RESULTS: We describe a generalization of the hierarchical clustering algorithm that efficiently incorporates these higher-order features by using a kernel function to map the data into a high-dimensional feature space. We then evaluate the utility of the kernel hierarchical clustering algorithm using both internal and external validation. The experiments demonstrate that the kernel representation itself is insufficient to provide improved clustering performance. We conclude that mapping gene expression data into a high-dimensional feature space is only a good idea when combined with a learning algorithm, such as the support vector machine that does not suffer from the curse of dimensionality. AVAILABILITY: Supplementary data at www.cs.columbia.edu/compbio/hiclust. Software source code available by request.

Algorithms↗

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations↗

Tree Killer, Qu'est-ce Que C'est? Insights From Forest Pathogen Genomes.

Forests are central to planetary health but are increasingly challenged by emerging diseases driven by climate change, global trade, and anthropogenic disturbance. Despite the apparent resilience of long-lived, genetically diverse tree hosts, forest ecosystems have repeatedly experienced landscape-level pathogen-driven transformations. Advances in genomics, transcriptomics, and functional biology have transformed our understanding of how fungal and oomycete pathogens interact with their hosts across a continuum of lifestyles, from saprotrophy and necrotrophy to biotrophy. Here, we synthesize insights from comparative and population genomics and functional studies across diverse forest pathosystems to examine the traits that characterize successful tree pathogens. We highlight how lifestyle plasticity, adaptations to woody tissues, vector-mediated transmission, and biotrophic stealth enable pathogens to colonize perennial hosts and persist over long temporal scales. We further examine how genome plasticity, hybridization, and horizontal gene transfer generate adaptive potential that often outpaces host evolutionary responses under current environmental change. Finally, we discuss emerging genomic tools, including biosurveillance, machine learning-based classification, and genome editing, that are beginning to link genotype to phenotype and inform assessments of disease risk. By integrating genomic, ecological, and evolutionary perspectives, this review outlines general principles governing forest pathogen success and identifies priorities for future research aimed at improving understanding, early detection, and management of forest diseases in a changing world.

Trees↗

Integrating single-cell transcriptomics to construct an oncogene-driven prognostic model and elucidate metabolic-immune crosstalk in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related deaths, its progression and treatment heterogeneity are mainly influenced by driver gene and tumor micro-environment (TME) interactions. Nevertheless, the mechanisms of this process at the single-cell level remain unclear. This study integrated TCGA and multi-center single-cell transcriptome data to identify a 575 genes HCC-specific core set, developing a single-cell "oncogene scoring" system to quantify individual carcinogenic activity. This score is significantly elevated in malignant and proliferative T cells and is closely associated with metabolic reprogramming, aberrant cell‒cell communication, and immunosuppressive phenotypes. Based on these characteristics, we constructed a machine learning-based Random Survival Forest (RSF) prognostic model validated in multiple independent cohorts, which classifies patients into distinct risk subtypes. The high-risk group exhibits genomic instability, increased tumor stemness, and immune evasion, while the low-risk group was more sensitive to drugs such as sorafenib. This study highlights the potential pathways by which high oncogenic activity is associated with HCC progression, suggesting a profound link with single-cell metabolic‒immune crosstalk. The constructed RSF model offers a promising computational framework for risk stratification and provides hypothesis-generating insights that may inform future personalized treatment strategies for HCC patients.

Hepatocellular carcinoma↗

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 μL of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell↗

PDP-Miner: an AI/ML tool to detect prophage tail proteins with depolymerase domains across thousands of bacterial genomes.

MOTIVATION: Antibiotic resistance is predicted to become the leading cause of human mortality by 2050. Despite this, no other major antibiotic class has been approved for medical use since 1987. Nevertheless, phage tail proteins offer a promising alternative, given their depolymerase activity toward outer membrane polysaccharides. Several pathogenic bacteria harbor prophages, thus making these prophages' molecular target already known. RESULTS: We therefore developed a wrapper for an existing machine learning-based phage depolymerase prediction tool (Depolymerase-Predictor), called PDP-Miner, which annotates phage tail proteins ab initio, detects depolymerase activity within this candidate protein subset, and then performs post-hoc validation by annotating protein domains thereby allowing the user to investigate for protein domains indicative of depolymerase activity. This tool allowed identification of 10 high confidence phage depolymerase gene candidates across all 1294 Pseudomonas genomes available on the International Pseudomonas Consortium Database while also accurately reporting depolymerases in known phage genomes, similarly to other software like PhageDPO or DepoScope. AVAILABILITY AND IMPLEMENTATION: Source code, test datasets and documentation are freely available for download at http:///www.github.com/jeffgauthier/pdpminer. This software is free and open source under the GNU General Public License v3.0.

Prophages↗

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies↗