Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

In silico analysis and comparison of the metabolic capabilities of different organisms by reducing metabolic complexity.

BACKGROUND: Understanding how metabolic capabilities diverge across microbial species is essential for deciphering community function, ecological interactions, and the design of synthetic microbiomes. Despite shared core pathways, microbial phenotypes can differ markedly due to evolutionary adaptations and metabolic specialization. Genome-scale metabolic models (GEMs) provide a systems-level framework to explore these differences; however, their complexity hinders direct comparison. RESULTS: We introduce NIS (Neidhardt-Ingraham-Schaechter), a computational workflow that integrates the redGEM, lumpGEM, and redGEMX algorithms to systematically reduce genome-scale models into biologically interpretable modules. This approach enables direct, quantitative comparison of fueling pathways, biomass biosynthetic routes, and environmental exchange processes while retaining essential metabolic information. We first demonstrate the utility of NIS by analyzing Escherichia coli and Saccharomyces cerevisiae, which revealed both conserved and divergent strategies in central metabolism, biosynthetic cost, and substrate utilization. We then applied NIS to the core honeybee gut microbiome, uncovering distinct metabolic traits, functional redundancy, and complementarity that help explain auxotrophy, cross-feeding interactions, and microbial coexistence. CONCLUSIONS: NIS provides an automated, scalable, and reproducible framework for dissecting microbial metabolic networks beyond gene content or taxonomy. By linking metabolism to ecological function, NIS offers new opportunities to interpret microbial community dynamics and to support the rational design of microbiomes in health, agriculture, and environmental applications. Video Abstract.

Metabolic Networks and Pathways

Decoding microbial metabolic complementarity from individual traits to community structuring.

A fundamental challenge in microbiome research lies in elucidating the functional capacity of microbial communities through community membership and genomic data. As community structuring and emergent functional traits are determined by bacterial community metabolic networks, it is important to gain insights into the principles that govern bacteria-bacteria interactions. Here, we applied an integrative framework linking individual strain-level traits to community structuring in a simplified synthetic bacterial community (SSC8) that promotes the growth of ungrafted watermelon. By combining mono- and coculture assays with genome-scale metabolic modeling and metabolomic profiling of spent media, we characterized directional interactions and resource dependencies among community members. Our findings show that positive interactions dominated the community network, accounting for 55% of all pairwise combinations, indicating a high prevalence of growth-promoting effects among strains. Genome-scale metabolic modeling showed that functional divergence among strains enhanced the potential for metabolic complementarity as phylogenetic distance increased. Integrating metabolic modeling with metabolomics further suggested that Pseudomonas azotifigens Q6 not only benefited from all other community members, but also exhibited mutualistic interactions with the other three strains, with metabolite exchange involving compounds such as L-lysine and L-cysteine. Pseudomonas azotifigens Q6 acted as an important driver of community composition by affecting the abundance of several other consortium members in vitro. These findings highlight the role of metabolic complementarity in driving community structuring by promoting selective persistence of specific strains. Our work provides mechanistic insights into microbial interaction networks in vitro and offers a conceptual foundation for the rational design of functionally robust and plant-beneficial microbiomes.

Bacteria

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans

HUMESS: integrating quantitative transcriptomic analysis and metabolic modeling to unveil condition-specific gene signatures.

SUMMARY: Transcriptomic analysis is a key tool for exploring gene expression, but the complexity of biological systems often limits its insights. In particular, the lack of intermodal or multi-layered analysis hinders the ability to fully capture key cellular functions such as metabolism from transcriptomic data alone. Here, we introduce a novel approach that informs transcriptomic data analysis with metabolic network modeling to address this. Unlike traditional methods, HUman MEtabolism Specific Signature (HUMESS) uses genome-scale metabolic modeling and flux analysis to highlight reactions and involved genes based on their metabolic significance, offering a deeper understanding of transcriptomic data. Our computational pipeline, supported by a user-friendly Rshiny application, enhances gene expression analysis by uncovering metabolic phenotypic signatures. AVAILABILITY AND IMPLEMENTATION: HUMESS is open source and available under GitLab https://gitlab.univ-nantes.fr/bird_pipeline_registry/humess with the complete documentation available at https://gitlab.univ-nantes.fr/bird_pipeline_registry/humess/-/wikis/Home. A zenodo archive is also available at the following DOI: https://doi.org/10.5281/zenodo.15487717. An RShiny application has been developed to facilitate the exploration and analysis of HUMESS's results. The app is available online at the following address: https://shiny-bird.univ-nantes.fr/app/shinymess but can also be installed locally, available under GitLab https://gitlab.univ-nantes.fr/pare-l/shinymess.

Humans

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning

A metabolic atlas of the Klebsiella pneumoniae species complex reveals lineage-specific metabolism and capacity for intra-species co-operation.

The Klebsiella pneumoniae species complex inhabits a wide variety of hosts and environments, and is a major cause of antimicrobial resistant infections. Genomics has revealed the population comprises multiple species/sub-species and hundreds of distinct co-circulating sub-lineage (SLs) that are associated with distinct gene complements. A substantial fraction of the pan-genome is predicted to be involved in metabolic functions and hence these data are consistent with metabolic differentiation at the SL level. However, this has so far remained unsubstantiated because in the past it was not possible to explore metabolic variation at scale. Here, we used a combination of comparative genomics and high-throughput genome-scale metabolic modeling to systematically explore metabolic diversity across the K. pneumoniae species complex (n = 7,835 genomes). We simulated growth outcomes for each isolate using carbon, nitrogen, phosphorus, and sulfur sources under aerobic and anaerobic conditions (n = 1,278 conditions per isolate). We showed that the distributions of metabolic genes and growth capabilities are structured in the population, and confirmed that SLs exhibit unique metabolic profiles. In vitro co-culture experiments demonstrated reciprocal commensalistic cross-feeding between SLs, effectively extending the range of conditions supporting individual growth. We propose that these substrate specializations may promote the existence and persistence of co-circulating SLs by reducing nutrient competition and facilitating commensal interactions. Our findings have implications for understanding the eco-evolutionary dynamics of K. pneumoniae and for the design of novel strategies to prevent opportunistic infections caused by this World Health Organization priority antimicrobial resistant pathogen.

Klebsiella pneumoniae

Metabolic network reconstruction as a resource for analyzing Salmonella Typhimurium SL1344 growth in the mouse intestine.

Nontyphoidal Salmonella strains (NTS) are among the most common foodborne enteropathogens and constitute a major cause of global morbidity and mortality, imposing a substantial burden on global health. The increasing antibiotic resistance of NTS bacteria has attracted a lot of research on understanding their modus operandi during infection. Growth in the gut lumen is a critical phase of the NTS infection. This might offer opportunities for intervention. However, the metabolic richness of the gut lumen environment and the inherent complexity and robustness of the metabolism of NTS bacteria call for modeling approaches to guide research efforts. In this study, we reconstructed a thermodynamically constrained and context-specific genome-scale metabolic model (GEM) for S. Typhimurium SL1344, a model strain well-studied in infection research. We combined sequence annotation, optimization methods and in vitro and in vivo experimental data. We used GEM to explore the nutritional requirements, the growth limiting metabolic genes, and the metabolic pathway usage of NTS bacteria in a rich environment simulating the murine gut. This work provides insight and hypotheses on the biochemical capabilities and requirements of SL1344 beyond the knowledge acquired through conventional sequence annotation and can inform future research aimed at better understanding NTS metabolism and identifying potential targets for infection prevention.

Salmonella typhimurium

Flux-sum coupling analysis of metabolic network models.

Metabolites acting as substrates and regulators of all biochemical reactions play an important role in maintaining the functionality of cellular metabolism. Despite advances in the constraint-based framework for genome-scale metabolic modeling, we lack reliable proxies for metabolite concentrations that can be efficiently determined and that allow us to investigate the relationship between metabolite concentrations in specific metabolic states in the absence of measurements. Here, we introduce a constraint-based approach, the flux-sum coupling analysis (FSCA), which facilitates the study of the interdependencies between metabolite concentrations by determining coupling relationships based on the flux-sum of metabolites. Application of FSCA on metabolic models of Escherichia coli, Saccharomyces cerevisiae, and Arabidopsis thaliana showed that the three coupling relationships are present in all models and pinpointed similarities in coupled metabolite pairs. Using the available concentration measurements of E. coli metabolites, we demonstrated that the coupling relationships identified by FSCA can capture the qualitative associations between metabolite concentrations and that flux-sum is a reliable proxy for metabolite concentration. Therefore, FSCA provides a novel tool for exploring and understanding the intricate interdependencies between the metabolite concentrations, advancing the understanding of metabolic regulation, and improving flux-centered systems biology approaches.

Escherichia coli

High-Quality Genome Assembly, Metabolome, Pangenome, and Metabolic Models of Megasphaera hexanoica KCCM 43214T.

Megasphaera hexanoica KCCM 43214T, isolated from cow rumen, is capable of producing medium-chain carboxylic acids such as hexanoate and octanoate. In this study, we present a high-quality genome assembly, along with intracellular metabolomic profiling and pangenomic analysis. Illumina sequencing generated 2.3 Gbp from 15,293,634 reads with a GC content of 49.5%, while PacBio HiFi sequencing produced 331.5 Mbp across 45,266 reads, with an average read length of 7,323 bp and a HiFi read N50 of 8,214 bp. Hybrid assembly of short and long reads resulted in a single 2.88 Mbp contig, containing 2,835 protein-coding genes. Genome-scale metabolic models were constructed to evaluate its metabolic capabilities under specific growth conditions. Intracellular metabolomic analysis of cells grown in medium containing fructose and lactate revealed key metabolic activities associated with chain elongation. Pangenomic analysis across nine annotated genomes identified 6,721 orthologous genes using OrthoMCL, emphasizing the genetic and functional diversity within the Megasphaera genus. This dataset offers valuable insights into the metabolism and biotechnological potential of M. hexanoica KCCM 43214T.

Metabolome

De Novo Genome Sequence Assembly of the Algal Endosymbiont Micractinium conductrix Derived From Its Host Paramecium bursaria 186b.

Endosymbiosis is a major driver of evolutionary innovation and underpins the function of diverse ecosystems. The origins and evolution of endosymbiosis are challenging to study experimentally due to the short-lived culturability of many microbial strains derived from endosymbiotic interactions. The facultative endosymbiosis between the ciliate, Paramecium bursaria, and the green alga, Micractinium conductrix (Chlorellaceae, Trebouxiophyceae), is ecologically widespread and has emerged as a powerful lab-tractable model system. This endosymbiosis is founded upon a reciprocal nutrient exchange, but each of the species can be cultured independently enabling quantification of symbiotic fitness effects, new partnerships to be generated in the lab, and co-associations to be subject to experimental evolution. To date, evolve-and-resequence approaches have been limited due to a lack of high-quality genome assemblies enabling gene variants to be identified. Here, we report a near telomere-to-telomere genome assembly for M. conductrix 186b, using a range of sequencing technologies. Comparative analysis shows that this is one of the most complete Chlorellaceae algal genome assemblies available to date. To aid accurate gene calling and annotation, we conducted both RNAseq and Iso-Seq transcriptome sequencing experiments. Collectively, these 'omics datasets will facilitate: (i) comparative genomics studies of endosymbiont evolution, (ii) evolve-and-resequence experiments, (iii) genome-scale metabolic modeling studies, and (iv) identification of targets for genetic modification experiments and biotechnological applications.

Symbiosis

Carbon metabolic homogenization is linked to microbial competition and antimicrobial resistance in soils under forest-to-cropland conversion.

Global agricultural expansion by converting natural forests into croplands often leads to soil functional homogenization and antimicrobial resistance enhancement, threatening ecosystem services. However, the associations between microbial carbon metabolic homogenization and antimicrobial resistance remain largely unknown. Here, we collected 240 paired forest and cropland soil samples from the most intensively farmed Yangtze River Basin in China, and constructed a novel framework based on microbial functional traits to decipher the role of carbon metabolic homogenization on antimicrobial resistance via microbial competition for metabolites. Using genome-scale metabolic models, we found that carbon metabolic homogenization was associated with a shift in microbial interactions from cooperation toward competition, with a 45.6% increase in competitive interactions that coincided with a 35.6% higher antimicrobial resistance gene (ARG) diversity. This shift was accompanied by smaller genome sizes and higher 16S rRNA copy numbers, indicating fast-growing, resource-acquisitive microbial strategies. Metabolic transfer analyses further revealed less cooperation relationships among microbial communities in cropland soils than in forest soils, indicating an intensified battle for communal metabolites and an attenuated exchange for complementary metabolites. Together, these findings provide a new framework to understand the association between carbon metabolic homogenization and soil antimicrobial resistance risks from the perspective of microbial traits and interactions under land use change.

Soil Microbiology

Function-based selection of synthetic communities enables mechanistic microbiome studies.

Understanding the complex interactions between microbes and their environment requires robust model systems such as synthetic communities (SynComs). We developed a functionally directed approach to generate SynComs by selecting strains that encode key functions identified in metagenomes. This approach enables the rapid construction of SynComs tailored to any ecosystem. To optimize community design, we implemented genome-scale metabolic models, providing in silico evidence for cooperative strain coexistence prior to experimental validation. Using this strategy, we designed multiple host-specific SynComs, including those for the rumen, mouse, and human microbiomes. By weighting functions differentially enriched in diseased versus healthy individuals, we constructed SynComs that capture complex host-microbe interactions. We designed an inflammatory bowel disease SynCom of 10 members that successfully induced colitis in gnotobiotic IL10-/- mice, demonstrating the potential of this method to model disease-associated microbiomes. Our study establishes a framework for designing functionally representative SynComs of any microbial ecosystem, facilitating mechanistic study.

Animals

Patient-specific modeling identifies metabolic interventions for reversing glucose use reprogramming in alcohol-associated hepatitis.

Alcoholic hepatitis (AH) is an acute form of alcohol-associated liver disease with very few treatment options. Recent studies highlighted liver metabolic reprogramming in AH as an indicator of severity. We aim at identifying new intervention points to reverse liver metabolic dysregulation across varying degrees of AH. We develop 89 personalized genome-scale metabolic models by integrating a generic human cellular metabolic model with liver transcriptomics data from AH patients with varying disease severity and healthy controls. We grade the AH patients based on the model-predicted level of glycolysis reprogramming and validate the results using published metabolomics data. We test in silico gene knockdown interventions to reverse the aberrant metabolic reprogramming in AH. Knockdown of two glycolytic genes, Hkdc1 and Pkm, significantly rebalance the metabolic fluxes toward a healthy liver metabolic phenotype. We use machine learning on the glycolysis fluxes to develop a quantitative glucose use reprogramming score, which correlates with AH severity and patient-specific responses to in silico gene knockdown interventions. The score was independently validated using a published AH liver transcriptomics dataset. We propose a cellular metabolism-based therapy targeting Hkdc1 and Pkm in the glycolysis pathway as a potential treatment for reversing the aberrant glucose metabolism in AH.

Humans

Membrane and proteome allocation constraints in Escherichia coli models during overflow metabolism.

The allocation of finite cellular resources is a fundamental principle that dictates microbial metabolic strategies and gives rise to complex phenomena, such as overflow metabolism, characterized by the production of respiro-fermentative by-products, including acetate, during rapid growth. Although proteome-constrained models have successfully predicted overflow metabolism in Escherichia coli, they often overlook the distinct biophysical and energetic costs associated with protein localization. The cellular membrane, in particular, represents a critical and constrained compartment where competition for space and synthesis machinery can create significant metabolic bottlenecks. To investigate this, we developed the membrane-associated constrained flux balance analysis (MAFBA), a scalable, genome-scale metabolic model that introduces a tunable constraint on the total protein mass allocated to the cellular membrane. Our model demonstrates that the overall and membrane-associated proteome allocation constraints interact to improve the accuracy of predicting the onset of overflow metabolism. It mechanistically reveals that at high growth rates, competition for limited membrane allocation forces a trade-off between growth-essential functions and respiratory capacity, leading to acetate production. Furthermore, MAFBA quantitatively explains the widely observed experimental phenomenon that expressing heterologous membrane proteins imposes a significantly higher metabolic burden than expressing cytosolic proteins. This study establishes membrane resource allocation as a key constraint governing bacterial physiology, acting in concert with overall proteome limitations. The resulting MAFBA framework provides a powerful and accessible tool for synthetic biology and metabolic engineering, enabling the prediction of metabolic costs associated with expressing membrane-bound proteins and guiding strain design strategies, holding promise for applications in bioproduction and metabolic engineering.

Escherichia coli

Reducing redundancy and enhancing accuracy through a phylogenetically-informed microbial community metabolic modeling approach.

MOTIVATION: Metabolic modeling has emerged as a powerful tool for predicting community functions. However, current modeling approaches face significant challenges in balancing the metabolic trade-offs between individual and community-level growth. In this study, we investigated the effect of metabolic relatedness among taxa on growth rate calculations by merging related taxa based on their metabolic similarity, introducing this approach as PhyloCOBRA. RESULTS: This approach enhanced the accuracy and efficiency of microbial community simulations by combining genome-scale metabolic models (GEMs) of closely related organisms, aligning with the concepts of niche differentiation and nestedness theory. To validate our approach, we implemented PhyloCOBRA within the MICOM and OptCom package (creating PhyloMICOM and PhyloOptCom, respectively), and applied it to metagenomic data from 186 individuals and four-species synthetic community (SynCom). Our results demonstrated significant improvement in the accuracy and reliability of growth rate predictions compared to the standard methods. Sensitivity analysis revealed that PhyloMICOM models were more robust to random noise, while Jaccard index calculations showed a reduction in redundancy, highlighting the enhanced specificity of the generated community models. Furthermore, PhyloMICOM reduced the computational complexity, addressing a key concern in microbial community simulations. This approach marks a significant advancement in community-scale metabolic modeling, offering a more stable, efficient, and ecologically relevant tool for simulating and understanding the intricate dynamics of microbial ecosystems. AVAILABILITY AND IMPLEMENTATION: PhyloCOBRA implementations are available as extensions to the MICOM packages and can be accessed at https://github.com/sepideh-mofidifar/PhyloCOBRA.

Phylogeny

Genomic features, metabolism, and biotechnological applications of Candida tropicalis and other non-albicans Candida species.

The production of bio-based products by yeasts from agroindustrial byproducts is a key strategy for advancing circular bioeconomy. While Saccharomyces species remain the predominant industrial yeasts, their limited ability to assimilate lactose, pentoses, and glycerol, as well as their sensitivity to lignocellulose-derived inhibitors, restricts their efficient application in bioprocesses based on using industrial byproducts as fermentation media. In contrast, several non-albicans Candida species exhibit broad substrate utilization capacities and enhanced tolerance to industrial stresses, making them attractive candidates for the bioconversion of agroindustrial residues. This review critically examines recent advances in the genomic, metabolic, and physiological characterization of promising non-albicans Candida species, including Candida tropicalis, Candida parapsilosis, Candida viswanathii, Candida sojae, and Candida maltosa. Emphasis is given to genome-scale metabolic models, carbon assimilation pathways, stress-response mechanisms, and metabolic engineering approaches aiming at the production of value-added compounds. By identifying current achievements, knowledge gaps, and biotechnological bottlenecks, this review highlights the potential of these yeasts as emerging platforms for sustainable bioprocesses within a circular bioeconomy framework.

Biotechnology

A global survey of taxa-metabolic associations across mouse microbiome communities.

Host-microbiota mutualism is rooted in the exchange of dietary and metabolic molecules. Microbial diversity broadens the metabolite pool, with each taxon contributing distinct compounds in varying proportions. In the human microbiome, high variability in consortial composition is largely compensated by similar metabolic functions across different taxa. However, the extent of compensation in lower diversity mouse models, and whether vivaria are metabolically equivalent, is unknown. We provide a searchable resource of microbiome composition variability across 51 murine vivaria and 12 wild mouse colonies worldwide, with vivarium-specific variants mapped according to predicted 3D structures for each microbial species. Our matched metabolomics data show that realized metabolic potential has relatively low variability, providing functional evidence for metabolic compensation. Additionally, variability is related to taxonomic composition rather than vivarium, revealing taxa-metabolite associations that are potentially relevant to phenotypic differences between vivaria. Collectively, this resource offers tools to strengthen microbiome studies and collaborative science.

Animals