Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Investigating metabolite essentiality through genome-scale analysis of Escherichia coli production capabilities.

MOTIVATION: A phenotype mechanism is classically derived through the study of a set of mutants and comparison of their biochemical capabilities. One method of comparing mutant capabilities is to characterize producible and knocked out metabolites. However such an effect is difficult to manually assess, especially for a large biochemical network and a complex media. Current algorithmic approaches towards analyzing metabolic networks either do not address this specific property or are computationally infeasible on the genome-scale. RESULTS: We have developed a novel genome-scale computational approach that identifies the full set of biochemical species that are knocked out from the metabolome following a gene deletion. Results from this approach are combined with data from in vivo mutant screens to examine the essentiality of metabolite production for a phenotype. This approach can also be a useful tool for metabolic network annotation validation and refinement in newly sequenced organisms. Combining an in silico genome-scale model of Escherichia coli metabolism with in vivo survival data, we uncover possible essential roles for several cell membranes, cell walls, and quinone species. We also identify specific biomass components whose production appears to be non-essential for survival, contrary to the assumptions of previous models. AVAILABILITY: Programs are available upon request from the authors in the form of Matlab script files. SUPPLEMENTARY INFORMATION: http://www.cis.upenn.edu/biocomp/manuscripts/bioinformatics_bti245/supp-info.html.

Algorithms↗

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning↗

In silico encounters: harnessing metabolic modelling to understand plant-microbe interactions.

Understanding plant-microbe interactions is vital for developing sustainable agricultural practices and mitigating the consequences of climate change on food security. Plant-microbe interactions can improve nutrient acquisition, reduce dependency on chemical fertilizers, affect plant health, growth, and yield, and impact plants' resistance to biotic and abiotic stresses. These interactions are largely driven by metabolic exchanges and can thus be understood through metabolic network modelling. Recent developments in genomics, metagenomics, phenotyping, and synthetic biology now enable researchers to harness the potential of metabolic modelling at the genome scale. Here, we review studies that utilize genome-scale metabolic modelling to study plant-microbe interactions in symbiotic, pathogenic, and microbial community systems. This review catalogues how metabolic modelling has advanced our understanding of the plant host and its associated microorganisms as a holobiont. We showcase how these models can contextualize heterogeneous datasets and serve as valuable tools to dissect and quantify underlying mechanisms. Finally, we consider studies that employ metabolic models as a testbed for in silico design of synthetic microbial communities with predefined traits. We conclude by discussing broader implications of the presented studies, future perspectives, and outstanding challenges.

Plants↗

Transcriptional regulation in constraints-based metabolic models of Escherichia coli.

Full genome sequences enable the construction of genome-scale in silico models of complex cellular functions. Genome-scale constraints-based models of Escherichia coli metabolism have been constructed and used to successfully interpret and predict cellular behavior under a range of conditions. These previous models do not account for regulation of gene transcription and thus cannot accurately predict some organism functions. Here we present an in silico model of the central E. coli metabolism that accounts for regulation of gene expression. This model accounts for 149 genes, the products of which include 16 regulatory proteins and 73 enzymes. These enzymes catalyze 113 reactions, 45 of which are controlled by transcriptional regulation. The combined metabolic/regulatory model can predict the ability of mutant E. coli strains to grow on defined media as well as time courses of cell growth, substrate uptake, metabolic by-product secretion, and qualitative gene expression under various conditions, as indicated by comparison with experimental data under a variety of environmental conditions. The in silico model may also be used to interpret dynamic behaviors observed in cell cultures. This combined metabolic/regulatory model is thus an important step toward the goal of synthesizing genome-scale models that accurately represent E. coli behavior.

Aerobiosis↗

Genome-scale microbial in silico models: the constraints-based approach.

Genome sequencing and annotation has enabled the reconstruction of genome-scale metabolic networks. The phenotypic functions that these networks allow for can be defined and studied using constraints-based models and in silico simulation. Several useful predictions have been obtained from such in silico models, including substrate preference, consequences of gene deletions, optimal growth patterns, outcomes of adaptive evolution and shifts in expression profiles. The success rate of these predictions is typically in the order of 70-90% depending on the organism studied and the type of prediction being made. These results are useful as a basis for iterative model building and for several practical applications.

Animals↗

Thermodynamics-based metabolic flux analysis.

A new form of metabolic flux analysis (MFA) called thermodynamics-based metabolic flux analysis (TMFA) is introduced with the capability of generating thermodynamically feasible flux and metabolite activity profiles on a genome scale. TMFA involves the use of a set of linear thermodynamic constraints in addition to the mass balance constraints typically used in MFA. TMFA produces flux distributions that do not contain any thermodynamically infeasible reactions or pathways, and it provides information about the free energy change of reactions and the range of metabolite activities in addition to reaction fluxes. TMFA is applied to study the thermodynamically feasible ranges for the fluxes and the Gibbs free energy change, Delta(r)G', of the reactions and the activities of the metabolites in the genome-scale metabolic model of Escherichia coli developed by Palsson and co-workers. In the TMFA of the genome scale model, the metabolite activities and reaction Delta(r)G' are able to achieve a wide range of values at optimal growth. The reaction dihydroorotase is identified as a possible thermodynamic bottleneck in E. coli metabolism with a Delta(r)G' constrained close to zero while numerous reactions are identified throughout metabolism for which Delta(r)G' is always highly negative regardless of metabolite concentrations. As it has been proposed previously, these reactions with exclusively negative Delta(r)G' might be candidates for cell regulation, and we find that a significant number of these reactions appear to be the first steps in the linear portion of numerous biosynthesis pathways. The thermodynamically feasible ranges for the concentration ratios ATP/ADP, NAD(P)/NAD(P)H, and H(extracellular)(+)/H(intracellular)(+) are also determined and found to encompass the values observed experimentally in every case. Further, we find that the NAD/NADH and NADP/NADPH ratios maintained in the cell are close to the minimum feasible ratio and maximum feasible ratio, respectively.

Escherichia coli↗

WILDkCAT: extract, retrieve, and predict enzyme turnover numbers of constraint-based metabolic models.

SUMMARY: Accurate enzyme turnover numbers are essential for building enzyme-constrained genome-scale metabolic models. However, collecting and curating these parameters remains a major bottleneck. Indeed, kcat values are scattered across multiple databases, reported under varying experimental conditions, and often missing for many enzymes. To address this challenge, we present WILDkCAT, a Python-based pipeline that enables the retrieval of kcat values from wild-type enzyme measured under user-specified pH and temperature ranges for a given metabolic model. The application to Escherichia coli (iML1515) and Homo sapiens (Human-GEM) models demonstrated the ability of WILDkCAT to retrieve substantial kcat coverage and its applicability across diverse genome-scale models. AVAILABILITY AND IMPLEMENTATION: WILDkCAT is available at https://github.com/sysbiolux/WILDkCAT and from PyPI. WILDkCAT works on all major operating systems and computer architectures. The documentation is available at https://sysbiolux.github.io/WILDkCAT.

Software↗

In vivo and in silico models of Drosophila for Parkinson's disease.

The fruit fly Drosophila melanogaster has emerged as an important model organism to shed light on neurodegeneration. Parkinson's disease (PD) is the second most prevalent neurodegenerative disorder, the cause of which is still mostly unclear. The long-term use of available PD drugs may have major side effects, and they only target the symptoms without providing any effective cure for the disease. Therefore, in vivo and in silico approaches are extensively used to model PD-like phenotypes in Drosophila and investigate cellular alterations underlying PD pathogenesis. In vivo models are particularly crucial to provide insight into the PD-related molecular processes. It has been a preferred approach to investigate these models by collecting omics datasets, which can be further analysed using in silico modeling such as genome-scale metabolic models and artificial intelligence applications. This review aims to summarise in vivo and in silico modeling studies in the literature to illustrate the potential of the Drosophila in the characterisation of PD-related biological mechanisms towards providing early biomarkers and novel treatment options for PD.

Humans↗

Evolutionary programming as a platform for in silico metabolic engineering.

BACKGROUND: Through genetic engineering it is possible to introduce targeted genetic changes and hereby engineer the metabolism of microbial cells with the objective to obtain desirable phenotypes. However, owing to the complexity of metabolic networks, both in terms of structure and regulation, it is often difficult to predict the effects of genetic modifications on the resulting phenotype. Recently genome-scale metabolic models have been compiled for several different microorganisms where structural and stoichiometric complexity is inherently accounted for. New algorithms are being developed by using genome-scale metabolic models that enable identification of gene knockout strategies for obtaining improved phenotypes. However, the problem of finding optimal gene deletion strategy is combinatorial and consequently the computational time increases exponentially with the size of the problem, and it is therefore interesting to develop new faster algorithms. RESULTS: In this study we report an evolutionary programming based method to rapidly identify gene deletion strategies for optimization of a desired phenotypic objective function. We illustrate the proposed method for two important design parameters in industrial fermentations, one linear and other non-linear, by using a genome-scale model of the yeast Saccharomyces cerevisiae. Potential metabolic engineering targets for improved production of succinic acid, glycerol and vanillin are identified and underlying flux changes for the predicted mutants are discussed. CONCLUSION: We show that evolutionary programming enables solving large gene knockout problems in relatively short computational time. The proposed algorithm also allows the optimization of non-linear objective functions or incorporation of non-linear constraints and additionally provides a family of close to optimal solutions. The identified metabolic engineering strategies suggest that non-intuitive genetic modifications span several different pathways and may be necessary for solving challenging metabolic engineering problems.

Algorithms↗

dAMN: a genome-scale neural-mechanistic hybrid model to predict bacterial growth dynamics.

SUMMARY: This study presents dAMN, a genome-scale neural-mechanistic hybrid model that combines neural networks with dynamic flux balance analysis to predict bacterial growth dynamics across diverse nutrient environments. Using a residual network architecture, dAMN predicts reaction fluxes and lag-phase parameters from initial medium composition, then integrates these predictions under stoichiometric constraints derived from genome-scale metabolic models. Trained on Escherichia coli and Pseudomonas putida growth datasets across combinatorial media, dAMN accurately forecasts temporal growth dynamics and generalizes to unseen media conditions, with mean R² ≥ 0.9. The model also reproduces biologically relevant behaviors including substrate depletion, acetate overflow, and diauxic shifts, while explicitly modeling lag phases usually absent from standard dFBA. AVAILABILITY AND IMPLEMENTATION: The dAMN software, associated models, and datasets are available at https://github.com/brsynth/dAMN-main-release and via Zenodo DOI: 10.5281/zenodo.17908125.

Escherichia coli↗

Systematic assignment of thermodynamic constraints in metabolic network models.

BACKGROUND: The availability of genome sequences for many organisms enabled the reconstruction of several genome-scale metabolic network models. Currently, significant efforts are put into the automated reconstruction of such models. For this, several computational tools have been developed that particularly assist in identifying and compiling the organism-specific lists of metabolic reactions. In contrast, the last step of the model reconstruction process, which is the definition of the thermodynamic constraints in terms of reaction directionalities, still needs to be done manually. No computational method exists that allows for an automated and systematic assignment of reaction directions in genome-scale models. RESULTS: We present an algorithm that - based on thermodynamics, network topology and heuristic rules - automatically assigns reaction directions in metabolic models such that the reaction network is thermodynamically feasible with respect to the production of energy equivalents. It first exploits all available experimentally derived Gibbs energies of formation to identify irreversible reactions. As these thermodynamic data are not available for all metabolites, in a next step, further reaction directions are assigned on the basis of network topology considerations and thermodynamics-based heuristic rules. Briefly, the algorithm identifies reaction subsets from the metabolic network that are able to convert low-energy co-substrates into their high-energy counterparts and thus net produce energy. Our algorithm aims at disabling such thermodynamically infeasible cyclic operation of reaction subnetworks by assigning reaction directions based on a set of thermodynamics-derived heuristic rules. We demonstrate our algorithm on a genome-scale metabolic model of E. coli. The introduced systematic direction assignment yielded 130 irreversible reactions (out of 920 total reactions), which corresponds to about 70% of all irreversible reactions that are required to disable thermodynamically infeasible energy production. CONCLUSION: Although not being fully comprehensive, our algorithm for systematic reaction direction assignment could define a significant number of irreversible reactions automatically with low computational effort. We envision that the presented algorithm is a valuable part of a computational framework that assists the automated reconstruction of genome-scale metabolic models.

Algorithms↗

In silico analysis and comparison of the metabolic capabilities of different organisms by reducing metabolic complexity.

BACKGROUND: Understanding how metabolic capabilities diverge across microbial species is essential for deciphering community function, ecological interactions, and the design of synthetic microbiomes. Despite shared core pathways, microbial phenotypes can differ markedly due to evolutionary adaptations and metabolic specialization. Genome-scale metabolic models (GEMs) provide a systems-level framework to explore these differences; however, their complexity hinders direct comparison. RESULTS: We introduce NIS (Neidhardt-Ingraham-Schaechter), a computational workflow that integrates the redGEM, lumpGEM, and redGEMX algorithms to systematically reduce genome-scale models into biologically interpretable modules. This approach enables direct, quantitative comparison of fueling pathways, biomass biosynthetic routes, and environmental exchange processes while retaining essential metabolic information. We first demonstrate the utility of NIS by analyzing Escherichia coli and Saccharomyces cerevisiae, which revealed both conserved and divergent strategies in central metabolism, biosynthetic cost, and substrate utilization. We then applied NIS to the core honeybee gut microbiome, uncovering distinct metabolic traits, functional redundancy, and complementarity that help explain auxotrophy, cross-feeding interactions, and microbial coexistence. CONCLUSIONS: NIS provides an automated, scalable, and reproducible framework for dissecting microbial metabolic networks beyond gene content or taxonomy. By linking metabolism to ecological function, NIS offers new opportunities to interpret microbial community dynamics and to support the rational design of microbiomes in health, agriculture, and environmental applications. Video Abstract.

Metabolic Networks and Pathways↗

Sequence-based analysis of metabolic demands for protein synthesis in prokaryotes.

Constraints-based models for microbial metabolism can currently be constructed on a genome-scale. These models do not account for RNA and protein synthesis. A scalable formalism to describe translation and transcription that can be integrated with the existing metabolic models is thus needed. Here, we developed such a formalism. The fundamental protein synthesis network described by this formalism was analysed via extreme pathway and flux balance analyses. The protein synthesis network exhibited one extreme pathway per messenger RNA synthesized and one extreme pathway per protein synthesized. The key parameters in this network included promoter strengths, messenger RNA half-lives, and the availability of nucleotide triphosphates, amino acids, RNA polymerase, and active ribosomes. Given these parameters, we were able to calculate a cell's material and energy expenditures for protein synthesis using a flux balance approach. The framework provided herein can subsequently be integrated with genome-scale metabolic models, providing a sequence-based accounting of the metabolic demands resulting from RNA and protein polymerization.

Amino Acids↗

Reconstruction of microbial transcriptional regulatory networks.

Although metabolic networks can be readily reconstructed through comparative genomics, the reconstruction of regulatory networks has been hindered by the relatively low level of evolutionary conservation of their molecular components. Recent developments in experimental techniques have allowed the generation of vast amounts of data related to regulatory networks. This data together with literature-derived knowledge has opened the way for genome-scale reconstruction of transcriptional regulatory networks. Large-scale regulatory network reconstructions can be converted to in silico models that allow systematic analysis of network behavior in response to changes in environmental conditions. These models can further be combined with genome-scale metabolic models to build integrated models of cellular function including both metabolism and its regulation.

Bacteria↗

Decoding microbial metabolic complementarity from individual traits to community structuring.

A fundamental challenge in microbiome research lies in elucidating the functional capacity of microbial communities through community membership and genomic data. As community structuring and emergent functional traits are determined by bacterial community metabolic networks, it is important to gain insights into the principles that govern bacteria-bacteria interactions. Here, we applied an integrative framework linking individual strain-level traits to community structuring in a simplified synthetic bacterial community (SSC8) that promotes the growth of ungrafted watermelon. By combining mono- and coculture assays with genome-scale metabolic modeling and metabolomic profiling of spent media, we characterized directional interactions and resource dependencies among community members. Our findings show that positive interactions dominated the community network, accounting for 55% of all pairwise combinations, indicating a high prevalence of growth-promoting effects among strains. Genome-scale metabolic modeling showed that functional divergence among strains enhanced the potential for metabolic complementarity as phylogenetic distance increased. Integrating metabolic modeling with metabolomics further suggested that Pseudomonas azotifigens Q6 not only benefited from all other community members, but also exhibited mutualistic interactions with the other three strains, with metabolite exchange involving compounds such as L-lysine and L-cysteine. Pseudomonas azotifigens Q6 acted as an important driver of community composition by affecting the abundance of several other consortium members in vitro. These findings highlight the role of metabolic complementarity in driving community structuring by promoting selective persistence of specific strains. Our work provides mechanistic insights into microbial interaction networks in vitro and offers a conceptual foundation for the rational design of functionally robust and plant-beneficial microbiomes.

Bacteria↗

Genome-scale thermodynamic analysis of Escherichia coli metabolism.

Genome-scale metabolic models are an invaluable tool for analyzing metabolic systems as they provide a more complete picture of the processes of metabolism. We have constructed a genome-scale metabolic model of Escherichia coli based on the iJR904 model developed by the Palsson Laboratory at the University of California at San Diego. Group contribution methods were utilized to estimate the standard Gibbs free energy change of every reaction in the constructed model. Reactions in the model were classified based on the activity of the reactions during optimal growth on glucose in aerobic media. The most thermodynamically unfavorable reactions involved in the production of biomass in E. coli were identified as ATP phosphoribosyltransferase, ATP synthase, methylene-tetra-hydrofolate dehydrogenase, and tryptophanase. The effect of a knockout of these reactions on the production of biomass and the production of individual biomass precursors was analyzed. Changes in the distribution of fluxes in the cell after knockout of these unfavorable reactions were also studied. The methodologies and results discussed can be used to facilitate the refinement of the feasible ranges for cellular parameters such as species concentrations and reaction rate constants.

ATP Phosphoribosyltransferase↗

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps↗

From genomes to in silico cells via metabolic networks.

Genome-scale metabolic models are the focal point of systems biology as they allow the collection of various data types in a form suitable for mathematical analysis. High-quality metabolic networks and metabolic networks with incorporated regulation have been successfully used for the analysis of phenotypes from phenotypic arrays and in gene-deletion studies. They have also been used for gene expression analysis guided by metabolic network structure, leading to the identification of commonly regulated genes. Thus, genome-scale metabolic modeling currently stands out as one of the most promising approaches to obtain an in silico prediction of cellular function based on the interaction of all of the cellular components.

Bacteria↗