Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome-scale reconstruction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans↗

Diet modulates cardiac metabolic stress during anthracycline treatment.

Diet is a modifiable determinant of cardiovascular risk and may influence tolerance to cancer therapies. The mechanisms by which specific dietary components affect cardiac metabolism during anthracycline treatment remain poorly defined, limiting the incorporation of dietary recommendations into treatment guidelines. Here, we integrated heart proteomics data from patients treated with or without anthracyclines with a genome-scale reconstruction of human cardiac metabolism (CardioNet). Using constraint-based flux analysis, we conducted >30,000 in silico simulations of diet scenarios generated from chemical profiles of ∼500 foods curated in the Periodic Table of Food Initiative. These simulations revealed that diets enriched in rapidly absorbable sugars and depleted of essential fatty acids impair cardiac metabolic efficiency, increasing reactive oxygen species production and the demand for purine salvage fluxes. These predicted metabolic patterns were consistent with plasma metabolomics from patients treated with anthracyclines, validating our findings. Computational modeling of 39 recipes across six cuisines revealed cardiometabolic effects of omnivorous versus vegan diets in patients. Modeling of a healthy vegan diet increased cardiometabolic efficiency compared with a healthy omnivorous diet in patients treated with anthracyclines, independent of the culinary background. Our approach demonstrates that integrating the molecular composition of food with genome-scale metabolic models enables systematic analysis of diet patterns for translational testing. Ultimately, these in silico studies provide a framework for trials and may inform dietary recommendations for improving cardiometabolic health.NEW & NOTEWORTHY We developed a systems biology framework to predict how diet influences cardiac metabolism during cancer therapy. Across >30,000 in silico diet simulations, we identified nutrient patterns that either exacerbate or mitigate anthracycline-induced metabolic stress. These findings demonstrate how computational modeling can uncover diet-metabolism interactions driving cardiotoxicity and guide dietary interventions.

Humans↗

RBC-GEM: A genome-scale metabolic model for systems biology of the human red blood cell.

Advancements with cost-effective, high-throughput omics technologies have had a transformative effect on both fundamental and translational research in the medical sciences. These advancements have facilitated a departure from the traditional view of human red blood cells (RBCs) as mere carriers of hemoglobin, devoid of significant biological complexity. Over the past decade, proteomic analyses have identified a growing number of different proteins present within RBCs, enabling systems biology analysis of their physiological functions. Here, we introduce RBC-GEM, one of the most comprehensive, curated genome-scale metabolic reconstructions of a specific human cell type to-date. It was developed through meta-analysis of proteomic data from 29 studies published over the past two decades resulting in an RBC proteome composed of more than 4,600 distinct proteins. Through workflow-guided manual curation, we have compiled the metabolic reactions carried out by this proteome to form a genome-scale metabolic model (GEM) of the RBC. RBC-GEM is hosted on a version-controlled GitHub repository, ensuring adherence to the standardized protocols for metabolic reconstruction quality control and data stewardship principles. RBC-GEM represents a metabolic network is a consisting of 820 genes encoding proteins acting on 1,685 unique metabolites through 2,723 biochemical reactions: a 740% size expansion over its predecessor. We demonstrated the utility of RBC-GEM by creating context-specific proteome-constrained models derived from proteomic data of stored RBCs for 616 blood donors, and classified reactions based on their simulated abundance dependence. This reconstruction as an up-to-date curated GEM can be used for contextualization of data and for the construction of a computational whole-cell models of the human RBC.

Humans↗

PMkbase (version 1.0): an interactive web-based tool for tracking bacterial metabolic traits using phenotype microarrays made interoperable with sequence information and visualizing/processing PM data.

Bacteria showcase remarkable metabolic diversity and traits, even among strains of the same species. In recent years, a large number of bacterial genomes have been sequenced, leading to the elucidation and documentation of genomic differences and commonalities across and within species. Genome-scale metabolic reconstructions, which are often defined and curated using data from phenotype microarrays, elucidate the differences in metabolic traits resulting from genomic diversity. These microarrays measure cellular respiration on a variety of carbon, nitrogen, phosphorus, and sulfur sources and various stressors and inhibitors over a period of time to determine the metabolic activity of a given strain. Despite their popularity in measuring bacterial metabolic activity and traits, no public databases that allow researchers to warehouse, access, and analyze this information currently exist. Additionally, there are no publicly available tools that allow researchers to view the variance of these metabolic traits across bacterial strains. To address this need, we present Phenotype Microarray Knowledgebase (PMkbase [version 1.0], https://pmkbase.com/), an interactive database that acts as a repository of phenotype microarray (PM) data with integrated sequence information. Binarized activity calls, along with associated kinetic parameters, are made for all metabolic substrates and inhibitors. Users can upload their own data for analysis and visualization and to perform quality checks on their experiments. PMkbase will address an unmet need to track and view bacterial metabolic traits and provide researchers with valuable information to develop metabolic models, enrich pangenomic analyses, and design new experiments.IMPORTANCEBacterial species can be differentiated by their metabolic profiles or the type of nutrients they consume. Interestingly, strains within the same species also display differences in nutrient consumption. Phenotype microarrays are a high-throughput, widely used technology to measure which substrates can be metabolized by various microbial strains and the extent to which inhibitors can affect it. Despite their widespread use, public databases to parse and access this data type at scale do not exist. PMkbase, which contains 9,024 data points for nitrogen substrate utilization, 41,664 data points for carbon substrate utilization, 8,448 data points for phosphorus/sulfur substrate utilization, and 27,264 data points on various antibiotics across three species (Escherichia coli, Pseudomonas putida, and Staphylococcus aureus), has been developed to allow researchers to freely access PM data, along with enriching the data with sequence information.

Bacteria↗

Genome-scale Gene Expression Analysis and Pathway Reconstruction in KEGG.

The massively parallel hybridization technologies by DNA chips and microarrays make it possible to monitor expression patterns of the whole set of genes in a genome under various conditions. The vast amount of data generated by such technologies necessitates the development of a new database management system that integrates expression data with other molecular biology databases and various analysis tools. We report here an extension of our KEGG (Kyoto Encyclopedia of Genes and Genomes) and DBGET/LinkDB systems for analyzing gene expression data in conjunction with pathway information and genomic information. It is now possible to make use of expression data for the reconstruction of pathways from the complete genome sequences.

Journal Article↗

Genome-scale metabolic modeling reveals increased reliance on valine catabolism in clinical isolates of Klebsiella pneumoniae.

Infections due to carbapenem-resistant Enterobacteriaceae have recently emerged as one of the most urgent threats to hospitalized patients within the United States and Europe. By far the most common etiological agent of these infections is Klebsiella pneumoniae, frequently manifesting in hospital-acquired pneumonia with a mortality rate of ~50% even with antimicrobial intervention. We performed transcriptomic analysis of data collected previously from in vitro characterization of both laboratory and clinical isolates which revealed shifts in expression of multiple master metabolic regulators across isolate types. Metabolism has been previously shown to be an effective target for antibacterial therapy, and genome-scale metabolic network reconstructions (GENREs) have provided a powerful means to accelerate identification of potential targets in silico. Combining these techniques with the transcriptome meta-analysis, we generated context-specific models of metabolism utilizing a well-curated GENRE of K. pneumoniae (iYL1228) to identify novel therapeutic targets. Functional metabolic analyses revealed that both composition and metabolic activity of clinical isolate-associated context-specific models significantly differs from laboratory isolate-associated models of the bacterium. Additionally, we identified increased catabolism of L-valine in clinical isolate-specific growth simulations. These findings warrant future studies for potential efficacy of valine transaminase inhibition as a target against K. pneumoniae infection.

Humans↗

Characterizing the metabolic phenotype: a phenotype phase plane analysis.

Genome-scale metabolic maps can be reconstructed from annotated genome sequence data, biochemical literature, bioinformatic analysis, and strain-specific information. Flux-balance analysis has been useful for qualitative and quantitative analysis of metabolic reconstructions. In the past, FBA has typically been performed in one growth condition at a time, thus giving a limited view of the metabolic capabilities of a metabolic network. We have broadened the use of FBA to map the optimal metabolic flux distribution onto a single plane, which is defined by the availability of two key substrates. A finite number of qualitatively distinct patterns of metabolic pathway utilization were identified in this plane, dividing it into discrete phases. The characteristics of these distinct phases are interpreted using ratios of shadow prices in the form of isoclines. The isoclines can be used to classify the state of the metabolic network. This methodology gives rise to a "phase plane" analysis of the metabolic genotype-phenotype relation relevant for a range of growth conditions. Phenotype phase planes (PhPPs) were generated for Escherichia coli growth on two carbon sources (acetate and glucose) at all levels of oxygenation, and the resulting optimal metabolic phenotypes were studied. Supplementary information can be downloaded from our website (http://epicurus.che.udel.edu).

Computational Biology↗

BioEMMA: Automated Generation of Model-Specific Escher-Compatible Maps from KEGG Pathways.

Genome-scale metabolic models are widely used to investigate cellular metabolism, but their interpretation and comparison are limited by the lack of reproducible pathway-level visualizations with a common spatial organization. This study presents BioEMMA, a Python-based tool for the automated generation of model-specific metabolic pathway maps in the Escher JSON format using coordinate information from curated KEGG pathway maps. BioEMMA parses KGML files, map reaction and metabolite identifiers to model database namespaces, filters pathway elements according to an input SBML model, adds non-primary metabolites, reconstructs Escher-compatible layouts, and supports flux visualization. The tool was integrated into a reproducible BioUML workflow for metabolic model reconstruction. BioEMMA was evaluated using the e_coli_core model and the KEGG glycolysis/gluconeogenesis pathway while generating a model-specific map with overlaid FBA fluxes. It was then applied to compare E. coli reconstructions generated by gapseq, ModelSEEDpy, and Reconstructor across three central carbon metabolism pathways. To broaden the evaluation, BioEMMA was applied using 87 prokaryotic BiGG models and three eukaryotic models. The analysis revealed pathway-specific differences in reaction coverage, shared and model-specific reactions, and predicted flux activity. BioEMMA therefore provides a reproducible framework for pathway-level visualization and comparison of genome-scale metabolic reconstructions within a common spatial coordinate system.

Escher maps↗

Regulation of gene expression in flux balance models of metabolism.

Genome-scale metabolic networks can now be reconstructed based on annotated genomic data augmented with biochemical and physiological information about the organism. Mathematical analysis can be performed to assess the capabilities of these reconstructed networks. The constraints-based framework, with flux balance analysis (FBA), has been used successfully to predict time course of growth and by-product secretion, effects of mutation and knock-outs, and gene expression profiles. However, FBA leads to incorrect predictions in situations where regulatory effects are a dominant influence on the behavior of the organism. Thus, there is a need to include regulatory events within FBA to broaden its scope and predictive capabilities. Here we represent transcriptional regulatory events as time-dependent constraints on the capabilities of a reconstructed metabolic network to further constrain the space of possible network functions. Using a simplified metabolic/regulatory network, growth is simulated under various conditions to illustrate systemic effects such as catabolite repression, the aerobic/anaerobic diauxic shift and amino acid biosynthesis pathway repression. The incorporation of transcriptional regulatory events in FBA enables us to interpret, analyse and predict the effects of transcriptional regulation on cellular metabolism at the systemic level.

Amino Acids↗

Metabolic network reconstruction as a resource for analyzing Salmonella Typhimurium SL1344 growth in the mouse intestine.

Nontyphoidal Salmonella strains (NTS) are among the most common foodborne enteropathogens and constitute a major cause of global morbidity and mortality, imposing a substantial burden on global health. The increasing antibiotic resistance of NTS bacteria has attracted a lot of research on understanding their modus operandi during infection. Growth in the gut lumen is a critical phase of the NTS infection. This might offer opportunities for intervention. However, the metabolic richness of the gut lumen environment and the inherent complexity and robustness of the metabolism of NTS bacteria call for modeling approaches to guide research efforts. In this study, we reconstructed a thermodynamically constrained and context-specific genome-scale metabolic model (GEM) for S. Typhimurium SL1344, a model strain well-studied in infection research. We combined sequence annotation, optimization methods and in vitro and in vivo experimental data. We used GEM to explore the nutritional requirements, the growth limiting metabolic genes, and the metabolic pathway usage of NTS bacteria in a rich environment simulating the murine gut. This work provides insight and hypotheses on the biochemical capabilities and requirements of SL1344 beyond the knowledge acquired through conventional sequence annotation and can inform future research aimed at better understanding NTS metabolism and identifying potential targets for infection prevention.

Salmonella typhimurium↗

Assessment of the metabolic capabilities of Haemophilus influenzae Rd through a genome-scale pathway analysis.

The annotated full DNA sequence is becoming available for a growing number of organisms. This information along with additional biochemical and strain-specific data can be used to define metabolic genotypes and reconstruct cellular metabolic networks. The first free-living organism for which the entire genomic sequence was established was Haemophilus influenzae. Its metabolic network is reconstructed herein and contains 461 reactions operating on 367 intracellular and 84 extracellular metabolites. With the metabolic reaction network established, it becomes necessary to determine its underlying pathway structure as defined by the set of extreme pathways. The H. influenzae metabolic network was subdivided into six subsystems and the extreme pathways determined for each subsystem based on stoichiometric, thermodynamic, and systems-specific constraints. Positive linear combinations of these pathways can be taken to determine the extreme pathways for the complete system. Since these pathways span the capabilities of the full system, they could be used to address a number of important physiological questions. First, they were used to reconcile and curate the sequence annotation by identifying reactions whose function was not supported in any of the extreme pathways. Second, they were used to predict gene products that should be co-regulated and perhaps co-expressed. Third, they were used to determine the composition of the minimal substrate requirements needed to support the production of 51 required metabolic products such as amino acids, nucleotides, phospholipids, etc. Fourth, sets of critical gene deletions from core metabolism were determined in the presence of the minimal substrate conditions and in more complete conditions reflecting the environmental niche of H. influenzae in the human host. In the former case, 11 genes were determined to be critical while six remained critical under the latter conditions. This study represents an important milestone in theoretical biology, namely the establishment of the first extreme pathway structure of a whole genome.

Genome, Bacterial↗

Understanding disease-associated metabolic changes in human colonic epithelial cells using the iColonEpithelium metabolic reconstruction.

The colonic epithelium plays a key role in the host-microbiome interactions, allowing uptake of various nutrients and driving important metabolic processes. To unravel detailed metabolic activities in the human colonic epithelium, our present study focuses on the generation of the first cell-type-specific genome-scale metabolic model (GEM) of human colonic epithelial cells, named iColonEpithelium. GEMs are powerful tools for exploring reactions and metabolites at the systems level and predicting the flux distributions at steady state. Our cell-type-specific iColonEpithelium metabolic reconstruction captures genes specifically expressed in the human colonic epithelial cells. iColonEpithelium is also capable of performing metabolic tasks specific to the colonic epithelium. A unique transport reaction compartment has been included to allow for the simulation of metabolic interactions with the gut microbiome. We used iColonEpithelium to identify metabolic signatures associated with inflammatory bowel disease. We used single-cell RNA sequencing data from Crohn's Diseases (CD) and ulcerative colitis (UC) samples to build disease-specific iColonEpithelium metabolic networks in order to predict metabolic signatures of colonocytes in both healthy and disease states. We identified reactions in nucleotide interconversion, fatty acid synthesis and tryptophan metabolism were differentially regulated in CD and UC conditions, relative to healthy control, which were in accordance with experimental results. The iColonEpithelium metabolic network can be used to identify mechanisms at the cellular level, and we show an initial proof-of-concept for how our tool can be leveraged to explore the metabolic interactions between host and gut microbiota.

Humans↗

Metabolic modeling of microbial strains in silico.

The large volume of genome-scale data that is being produced and made available in databases on the World Wide Web is demanding the development of integrated mathematical models of cellular processes. The analysis of reconstructed metabolic networks as systems leads to the development of an in silico or computer representation of collections of cellular metabolic constituents, their interactions and their integrated function as a whole. The use of quantitative analysis methods to generate testable hypotheses and drive experimentation at a whole-genome level signals the advent of a systemic modeling approach to cellular and molecular biology.

Genome↗

Can't see the forest for the trees: The influence of marker type on inferred phylogenetic relationships in a cosmopolitan bat genus.

Fine-resolution information on species relationships and biological diversity is critically needed to guide conservation efforts amidst rapid environmental changes. Systematics, which forms the foundation of this knowledge, has been revolutionized by phylogenomics, utilizing genome-scale datasets. However, the use of diverse marker types, non-comparable taxon sampling, and outgroup selection can lead to conflicting phylogenetic hypotheses. These inconsistencies complicate study comparisons and hinder our ability to assess marker-specific impacts on phylogenetic resolution. The phylogenetic reconstruction of the bat genus Myotis, encompassing over 140 species and characterized by a rapid radiation in the last 20 million years, has been particularly influenced by these challenges. Achieving phylogenetic resolution in Myotis is particularly complex due to subtle interspecific differences in both morphological and molecular traits. Mitochondrial and nuclear markers often produce discordant trees, influenced by hybridization, introgression, and methodological variations. In this study, we employed a consistent taxonomic sample set of 44 Myotis taxa to evaluate the impact of five different genetic marker types on phylogenetic reconstruction. We observed significant discordance between topologies derived from conserved nuclear and mitochondrial markers and found that transposable elements were inadequate for resolving relationships across the entire genus. Our results also clarify the placement of previously problematic taxa within the genus. These findings emphasize the importance of aligning genetic marker choice with specific phylogenetic questions and highlight the influence of taxonomic and methodological variation on phylogenomic outcomes. This work provides a framework for improving phylogenetic inference in rapidly radiating groups and enhances our understanding of evolutionary history in Myotis.

Animals↗

EcoGene: a genome sequence database for Escherichia coli K-12.

The EcoGene database provides a set of gene and protein sequences derived from the genome sequence of Escherichia coli K-12. EcoGene is a source of re-annotated sequences for the SWISS-PROT and Colibri databases. EcoGene is used for genetic and physical map compilations in collaboration with the Coli Genetic Stock Center. The EcoGene12 release includes 4293 genes. EcoGene12 differs from the GenBank annotation of the complete genome sequence in several ways, including (i) the revision of 706 predicted or confirmed gene start sites, (ii) the correction or hypothetical reconstruction of 61 frame-shifts caused by either sequence error or mutation, (iii) the reconstruction of 14 protein sequences interrupted by the insertion of IS elements, and (iv) pre-dictions that 92 genes are partially deleted gene fragments. A literature survey identified 717 proteins whose N-terminal amino acids have been verified by sequencing. 12 446 cross-references to 6835 literature citations and s are provided. EcoGene is accessible at a new website: http://bmb.med.miami.edu/EcoGene/EcoWeb. Users can search and retrieve individual EcoGene GenePages or they can download large datasets for incorporation into database management systems, facilitating various genome-scale computational and functional analyses.

Databases, Factual↗

Network methods for diagonal integration of unpaired single-cell multiomics data: a review.

MOTIVATION: Advances in single-cell sequencing have enabled multiomics profiling at unprecedented resolution; however, mass spectrometry-based single-cell proteomics (scMS) remains inherently destructive, precluding simultaneous transcriptomic capture. Unlike antibody-based methods such as CITE-seq, which permit paired profiling but are restricted to targeted protein panels, scMS provides unbiased, genome-scale coverage of the intracellular proteome yet necessitates post hoc integration of unpaired datasets. This diagonal integration challenge, where transcriptomes and proteomes are measured in separate cells lacking shared anchors, remains underserved by existing reviews, which focus predominantly on vertical integration strategies enabled by non-destructive assays. RESULTS: We survey the complete computational pipeline for constructing mechanistic proteogenomic networks from unpaired single-cell data, covering: (i) unimodal network inference such as knowledge-based approaches, probabilistic graphical models, temporal directionality inference, and generative and foundation model strategies that establish the transcriptomic scaffold; (ii) cross-modal integration architectures such as network propagation, graph neural networks (scMRDR, scmFormer, scCotag), and consensus frameworks designed explicitly for the unpaired proteomics setting; and (iii) benchmarking paradigms spanning network reconstruction (BEELINE, GRETA, CausalBench) and multi-task integration evaluation (scMultiBench, SCMMIB), with guidance on metric selection under network sparsity and class imbalance. We identify three principal axes of future development: generative proteomic translation from transcriptomic precursors, inductive prior embedding in next-generation architectures, and perturbation-based causal benchmarking. AVAILABILITY AND IMPLEMENTATION: This is a review article; no novel software is distributed. A curated benchmark resource table, methods starter guide, and per-method bottleneck annotations are provided in the Supplementary Material.

Multiomics↗

Metabolism and gene expression models for the microbiome reveal how diet and metabolic dysbiosis impact disease.

The gut microbiome plays a critical role in human health, spurring extensive research using multi-omic technologies. Although these tools offer valuable insights, they often fall short in capturing the complexity of microbial interactions that associate with disease onset, progression, and treatment. Thus, integration of multi-omics datasets with metabolic models is needed to predict associations between microbial activity and disease. Here, we automated the reconstruction of 495 metabolic and gene expression models (ME-models), overcoming the main limitation preventing the wide use of this approach. We integrated them with multi-omics data from patients with inflammatory bowel disease (IBD), identifying taxa associated with variations in amino acids, short-chain fatty acids, and pH in the gut of IBD patients. In general, this approach provides testable hypotheses of the metabolic activity of the gut microbiota, and the automated pipeline opens the opportunity to study microbial interactions in other biologically relevant settings using ME-models.

Humans↗

Genomic insights into natural selection in recent human history.

For over a century, scientists have debated the extent to which genetic and phenotypic variation among present-day humans is the result of natural selection - in which heritable traits influence survival or reproduction - versus neutral processes such as genetic drift or population history. The initial sequencing of the human genome and subsequent population resequencing studies enabled genome-scale searches for signatures of selection in present-day genomes. This first generation of genome-wide selection scans identified many targets but left open questions about the timing and nature of selection, making it challenging to identify environmental and biological drivers. Recent methodological advances based on reconstructing ancestral recombination graphs have increased the potential power and resolution of selection scans based on present-day genomes, while the availability of new data on ancient DNA has facilitated the direct reconstruction of genetic change through time. However, there is little consensus on how to use these data to detect and interpret signatures of selection, while avoiding confounders. Here, we review the current state of knowledge about the impact of selection on human genomic diversity and highlight conceptual advances in our understanding of human evolution over the past 10,000 years.

Journal Article↗