Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

High-throughput classification of yeast mutants for functional genomics using metabolic footprinting.

Many technologies have been developed to help explain the function of genes discovered by systematic genome sequencing. At present, transcriptome and proteome studies dominate large-scale functional analysis strategies. Yet the metabolome, because it is 'downstream', should show greater effects of genetic or physiological changes and thus should be much closer to the phenotype of the organism. We earlier presented a functional analysis strategy that used metabolic fingerprinting to reveal the phenotype of silent mutations of yeast genes. However, this is difficult to scale up for high-throughput screening. Here we present an alternative that has the required throughput (2 min per sample). This 'metabolic footprinting' approach recognizes the significance of 'overflow metabolism' in appropriate media. Measuring intracellular metabolites is time-consuming and subject to technical difficulties caused by the rapid turnover of intracellular metabolites and the need to quench metabolism and separate metabolites from the extracellular space. We therefore focused instead on direct, noninvasive, mass spectrometric monitoring of extracellular metabolites in spent culture medium. Metabolic footprinting can distinguish between different physiological states of wild-type yeast and between yeast single-gene deletion mutants even from related areas of metabolism. By using appropriate clustering and machine learning techniques, the latter based on genetic programming, we show that metabolic footprinting is an effective method to classify 'unknown' mutants by genetic defect.

Cells, Cultured↗

Social disconnection integrates genetic and proteomic risks in suicidal ideation and depression.

Suicidal ideation (SI) and major depressive disorder (MDD) are complex psychiatric conditions arising from the interplay of genetic liability, molecular processes, and psychosocial factors. While these dimensions have been extensively studied in isolation, their joint contribution to SI and MDD remains unclear. This study integrates multi-modal data to elucidate these synergistic effects and develop robust models for individual-level risk stratification. Leveraging longitudinal multi-modal data from 13,085 UK Biobank participants, we integrated genomic, proteomic, and social connection profiles. We developed interpretable risk scores using a rigorous supervised machine learning framework encompassing diverse linear and ensemble classifiers. Permutation importance was employed to quantify feature contributions and derive transparent, weighted risk metrics across diverse classifiers. These scores were validated through association, interaction, and mediation analyses. Social connection-based risk scores significantly differentiated cases and controls across the two suicidal ideation phenotypes at 2017 and 2023 with cross-sectional analyses (AUCs: 0.70 - 0.73), outperforming proteomic-only models. Functional dimensions of social connection emerged as the most informative predictors. Longitudinal analyses revealed that social risk scores at baseline predicted suicidal ideation onset six years later, independent of demographic covariates. Interaction analyses demonstrated that polygenic risk for suicide attempt significantly interacted with both social and proteomic risk features in relation to depression. Structural equation models further confirmed that social disconnection acts as a key mediator linking genetic predisposition to MDD and SI. Social disconnection is a critical risk factor mediating the impact of genetic vulnerability on psychiatric outcomes. Integrating social, genetic, and molecular data supports a multilevel framework for risk stratification and highlights the potential of socially oriented interventions to mitigate biological risk.

Humans↗

Microbiome features associated with persistent intestinal carriages of Escherichia coli ST131 in a Southeast Asian cohort study.

Escherichia coli sequence-type 131 (ST131) is the dominant global extraintestinal pathogen capable of asymptomatic intestinal carriage and sustained household transmission, challenging infection control. Despite its clinical significance, the ecological determinants of gut persistence remain poorly understood. We performed shotgun metagenomics on fecal samples to investigate gut microbiome features associated with ST131-positive samples, distinct host carrier statuses (persistent, intermittent and non-carriers) and household risks in a study of a Southeast Asian cohort. Here, we show that ST131 carriage was associated with compositional shifts without reducing species alpha-diversity. Regression analyses identified depletion of commensal taxa and the 1,5-anhydrofructose degradation pathway in ST131-positive samples. Persistent carriers exhibited highly perturbed microbiome enriched with pathobionts, aerobactin- and lipopolysaccharide (LPS)-biosynthesis pathways. Comparing household risk groups to control, revealed that biotin biosynthesis and 1,5-anhydrofructose degradation may influence ST131 co-colonization through both direct and indirect mechanisms. Machine learning analyses identified metabolic pathways as stronger discriminators of persistent carriage than taxonomic features. Genomic-resolved analysis of clinical ST131 isolates revealed conserved genes for iron-acquisition, LPS and antibiotic resistance determinants. Overall, while commensals and metabolism may influence initial ST131 colonization, persistent carriage is associated with specific microbial and metabolic adaptations, providing potential targets to limit intestinal ST131 persistence.

Humans↗

Proteo-metabolomic integration identifies stage-specific candidate biomarkers for Parkinson's disease.

Parkinson's disease (PD) is a progressive neurodegenerative disorder with a prolonged prodromal phase and complex motor symptoms. Despite improved clinical criteria, early diagnosis and longitudinal monitoring remain challenging. While cerebrospinal fluid (CSF) and plasma metabolites and proteins show biomarker potential, their utility in predictive models is insufficiently characterized. We employed a secondary computational approach to integrate proteometabolomic profiles from CSF and plasma samples of >1100 Parkinson's Progression Markers Initiative (PPMI) participants. Using multi-omics machine learning, we identified biofluid-specific signatures and evaluated predictive performance. Twenty-one biomarker candidates were validated across three models (SVM, GLMNET, RF); SVM and GLMNET achieved the highest recall (83-86%) and AUCs of 0.84-0.89. Longitudinal mixed-effects modeling revealed eight candidates associated with progression across diagnostic stages. We identified a three-part molecular framework characterizing neurodegeneration: a diagnostic subpanel reflecting early microbiome dysregulation (secretory granins and metabolites) and synaptic breakdown; a second subpanel monitoring phenoconversion via neurogenesis precursors and extracellular matrix proteins; and a third subpanel tracking progression through chronic neuroinflammation and immune activation. This integrated multi-omics approach provides a robust framework for stage-specific PD monitoring and potential clinical deployment.

Journal Article↗

Global microbial DNA signatures of temperature and nutrient limitation across ecosystems.

Microbial genomes continuously adapt to environmental conditions, but identifying universal signatures of adaptation remains challenging. Here we show that environmental temperature can be accurately predicted across ecosystems from DNA composition alone (R2 = 0.75), using tetranucleotide frequencies from 1,235 marine and soil metagenomes and a machine learning approach. This predictive signal was also apparent within individual taxa, consistent with a fundamental temperature-associated signature. By contrast, GC content exhibited opposite correlations with temperature in soil (positive) and marine (negative) environments. This phenomenon was probably driven by differences in nutrient availability, as GC content increases with nutrients while nutrients decrease with temperature in marine samples. By integrating these observations, we identified specific tetranucleotides, with 50% GC, that displayed consistent and robust temperature correlations across environments and may have contributed to the stability of predictions. This work highlights metagenome-wide DNA-temperature associations, relevant for understanding microbial community responses to global changes.

Journal Article↗

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans↗

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article↗

A genomic catalog of Earth's bacterial and archaeal symbionts.

Microbial symbiosis drives the functional and phylogenomic diversification of life on Earth yet remains underexplored because of culturing challenges. This study used machine learning (ML) to predict symbiotic lifestyles in more than a hundred thousand microbial genomes from diverse environmental metagenome samples and reference genomes. Predictions were performed using symclatron, an ML framework developed to identify genomic signatures of symbionts. Predictions were deposited in a catalog we established called Symbiont Genomes (SymGs). The results indicate that 15-23% of uncultivated microorganisms likely engage in symbiotic relationships with other organisms, categorized as host-associated or obligate intracellular lifestyles, and are present in half of all known bacterial and archaeal phyla. We also identify genomic signatures of symbiotic lifestyles, including the loss of certain metabolic functions and the differential presence of metabolic modules that may enable host-dependent living. The symclatron software and the SymGs catalog represent valuable resources for studying symbioses, potentially facilitating future mechanistic investigations and engineering of host-microorganism associations.

Journal Article↗

Mechanisms underlying disease-causing variants in promoters and enhancers.

The study of human monogenic disorders has been a powerful tool for generating a deep understanding of protein function/dysfunction and for uncovering underlying biological mechanisms. Here we explore the insights that an expanding catalog of noncoding monogenic disease variants can provide into the functions of the noncoding genome. We focus on small genetic alterations (one to a few tens of base pairs) in cis-regulatory elements-promoters, enhancers and silencers-and their potential mechanisms of action, such as loss or gain of function. We discuss the challenges in determining pathogenicity for variants in the noncoding genome, discuss why there might be so few concrete examples and highlight the opportunities for advancing this area of human genetics by using experimental and machine-learning tools.

Journal Article↗

Patient-specific modeling identifies metabolic interventions for reversing glucose use reprogramming in alcohol-associated hepatitis.

Alcoholic hepatitis (AH) is an acute form of alcohol-associated liver disease with very few treatment options. Recent studies highlighted liver metabolic reprogramming in AH as an indicator of severity. We aim at identifying new intervention points to reverse liver metabolic dysregulation across varying degrees of AH. We develop 89 personalized genome-scale metabolic models by integrating a generic human cellular metabolic model with liver transcriptomics data from AH patients with varying disease severity and healthy controls. We grade the AH patients based on the model-predicted level of glycolysis reprogramming and validate the results using published metabolomics data. We test in silico gene knockdown interventions to reverse the aberrant metabolic reprogramming in AH. Knockdown of two glycolytic genes, Hkdc1 and Pkm, significantly rebalance the metabolic fluxes toward a healthy liver metabolic phenotype. We use machine learning on the glycolysis fluxes to develop a quantitative glucose use reprogramming score, which correlates with AH severity and patient-specific responses to in silico gene knockdown interventions. The score was independently validated using a published AH liver transcriptomics dataset. We propose a cellular metabolism-based therapy targeting Hkdc1 and Pkm in the glycolysis pathway as a potential treatment for reversing the aberrant glucose metabolism in AH.

Humans↗

Human-centric intelligent systems for exploration and knowledge discovery.

This speculative article discusses research and development relating to computational intelligence (CI) technologies comprising powerful machine-based search and exploration techniques that can generate, extract, process and present high-quality information from complex, poorly understood biotechnology domains. The integration and capture of user experiential knowledge within such CI systems in order to support and stimulate knowledge discovery and increase scientific and technological understanding is of particular interest. The manner in which appropriate user interaction can overcome problems relating to poor problem representation within systems utilising evolutionary computation (EC), machine-learning and software agent technologies is investigated. The objective is the development of user-centric intelligent systems that support an improving knowledge-base founded upon gradual problem re-definition and reformulation. Such an approach can overcome initial lack of understanding and associated uncertainty.

Artificial Intelligence↗

From sequence space to ecological function: microbiome-derived antimicrobial peptides as community effectors and therapeutic leads.

Antimicrobial peptide research has long centred on host defence molecules, yet microbiomes themselves encode a diverse and increasingly important repertoire of peptide-based antimicrobials. These microbiome-derived antimicrobial peptides include bacteriocins, ribosomally synthesised and post-translationally modified peptides, cryptic short open reading frame-encoded peptides, embedded antimicrobial regions within larger proteins, and selected peptide antibiotics recovered from human, animal, plant and environmental microbiomes. Recent advances in genome mining, metagenomics, and machine learning have greatly expanded the scale of discovery, moving the field from a handful of landmark exemplars to large candidate catalogues spanning the global microbiome. In the clearest cases, these molecules are not only anti-infective leads but ecological effectors: they mediate microbial competition, enforce colonisation resistance, and influence community structure within densely occupied niches. The present review synthesises the field across discovery classes, microbiome sources, ecological roles, and translational bottlenecks, emphasizing a central limitation of the field: candidate catalogues are expanding at extraordinary scale, while evidence for native expression, producer assignment, ecological function, and in vivo relevance remains limited for the vast majority of predicted molecules. Progress will depend on workflows that connect sequence level prediction to biological context through expression support, producer assignment, community level validation, and perturbation-based approaches that distinguish ecological association from causal function. Microbiome-derived antimicrobial peptides are best understood not only as promising therapeutic leads, but also as molecular mediators of microbial social life whose ecological origins are central to their interpretation and future application.

Microbiota↗

Toward diagnostic and phenotype markers for genetically transmitted speech delay.

Converging evidence supports the hypothesis that the most common subtype of childhood speech sound disorder (SSD) of currently unknown origin is genetically transmitted. We report the first findings toward a set of diagnostic markers to differentiate this proposed etiological subtype (provisionally termed speech delay-genetic) from other proposed subtypes of SSD of unknown origin. Conversational speech samples from 72 preschool children with speech delay of unknown origin from 3 research centers were selected from an audio archive. Participants differed on the number of biological, nuclear family members (0 or 2+) classified as positive for current and/or prior speech-language disorder. Although participants in the 2 groups were found to have similar speech competence, as indexed by their Percentage of Consonants Correct scores, their speech error patterns differed significantly in 3 ways. Compared with children who may have reduced genetic load for speech delay (no affected nuclear family members), children with possibly higher genetic load (2+ affected members) had (a) a significantly higher proportion of relative omission errors on the Late-8 consonants; (b) a significantly lower proportion of relative distortion errors on these consonants, particularly on the sibilant fricatives /s/, /z/, and //; and (c) a significantly lower proportion of backed /s/ distortions, as assessed by both perceptual and acoustic methods. Machine learning routines identified a 3-part classification rule that included differential weightings of these variables. The classification rule had diagnostic accuracy value of 0.83 (95% confidence limits = 0.74-0.92), with positive and negative likelihood ratios of 9.6 (95% confidence limits = 3.1-29.9) and 0.40 (95% confidence limits = 0.24-0.68), respectively. The diagnostic accuracy findings are viewed as promising. The error pattern for this proposed subtype of SSD is viewed as consistent with the cognitive-linguistic processing deficits that have been reported for genetically transmitted verbal disorders.

Child, Preschool↗

Combination of multimodal imaging and molecular genetic information to investigate complex psychiatric disorders.

Multimodal imaging, the combination of several brain imaging techniques in one subject, provides a wealth of parameters and favours the interpretation of complex models in schizophrenia research. Moreover, new imaging tools allow the investigation of distinct neurotransmitter systems and their modulation by pharmacological intervention. An important feature of multimodal imaging is the possibility to characterize the activation dependencies of different neurotransmitters and provide the experimental tool to test system models of brain function and dysfunction. The combination of measurement techniques with high temporal resolution (e. g. MEG, EEG) and high spatial resolution (e. g. fMRI) facilitate the understanding of local and global systems as well as time characteristics. Moreover, the association of imaging parameters with genetic variations of neurotransmitter systems allows the investigation of neurotransmitter activity and its role in the pathophysiology of schizophrenia. To overcome the limitations of standard statistical methods, new approaches in machine learning have to be adapted to handle multiple parameters obtained from brain imaging and genetic measurements.

Brain↗

Error criteria for cross validation in the context of chaotic time series prediction.

The prediction of a chaotic time series over a long horizon is commonly done by iterating one-step-ahead prediction. Prediction can be implemented using machine learning methods, such as radial basis function networks. Typically, cross validation is used to select prediction models based on mean squared error. The bias-variance dilemma dictates that there is an inevitable tradeoff between bias and variance. However, invariants of chaotic systems are unchanged by linear transformations; thus, the bias component may be irrelevant to model selection in the context of chaotic time series prediction. Hence, the use of error variance for model selection, instead of mean squared error, is examined. Clipping is introduced, as a simple way to stabilize iterated predictions. It is shown that using the error variance for model selection, in combination with clipping, may result in better models.

Journal Article↗

Finding scientific topics.

A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.

Databases, Factual↗

Mapping subsets of scholarly information.

We illustrate the use of machine learning techniques to analyze, structure, maintain, and evolve a large online corpus of academic literature. An emerging field of research can be identified as part of an existing corpus, permitting the implementation of a more coherent community structure for its practitioners.

Computers↗

Inferring network mechanisms: the Drosophila melanogaster protein interaction network.

Naturally occurring networks exhibit quantitative features revealing underlying growth mechanisms. Numerous network mechanisms have recently been proposed to reproduce specific properties such as degree distributions or clustering coefficients. We present a method for inferring the mechanism most accurately capturing a given network topology, exploiting discriminative tools from machine learning. The Drosophila melanogaster protein network is confidently and robustly (to noise and training data subsampling) classified as a duplication-mutation-complementation network over preferential attachment, small-world, and a duplication-mutation mechanism without complementation. Systematic classification, rather than statistical study of specific properties, provides a discriminative approach to understand the design of complex networks.

Algorithms↗