Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Modelling strategies for the industrial exploitation of lactic acid bacteria.

Lactic acid bacteria (LAB) have a long tradition of use in the food industry, and the number and diversity of their applications has increased considerably over the years. Traditionally, process optimization for these applications involved both strain selection and trial and error. More recently, metabolic engineering has emerged as a discipline that focuses on the rational improvement of industrially useful strains. In the post-genomic era, metabolic engineering increasingly benefits from systems biology, an approach that combines mathematical modelling techniques with functional-genomics data to build models for biological interpretation and--ultimately--prediction. In this review, the industrial applications of LAB are mapped onto available global, genome-scale metabolic modelling techniques to evaluate the extent to which functional genomics and systems biology can live up to their industrial promise.

Biomass↗

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans↗

Automated protein modelling--the proteome in 3D.

Functional analysis of the proteins discovered in fully sequenced genomes represent the next major challenge of life science research. Computational methods play an increasingly important role in this activity. Among them, comparative protein modelling will play a major role in this challenge, especially in the light of the Structural Genomics programmes about to be started around the world. In recent years, much progress has been made in automating these methods, enabling the production of models for genome scale problems. In this review we discuss how protein models can be applied to functional analysis, as well as some of the current issues and limitations inherent to these methods.

Animals↗

In silico aided metabolic engineering of Saccharomyces cerevisiae for improved bioethanol production.

In silico genome-scale cell models are promising tools for accelerating the design of cells with improved and desired properties. We demonstrated this by using a genome-scale reconstructed metabolic network of Saccharomyces cerevisiae to score a number of strategies for metabolic engineering of the redox metabolism that will lead to decreased glycerol and increased ethanol yields on glucose under anaerobic conditions. The best-scored strategies were predicted to completely eliminate formation of glycerol and increase ethanol yield with 10%. We successfully pursued one of the best strategies by expressing a non-phosphorylating, NADP(+)-dependent glyceraldehyde-3-phosphate dehydrogenase in S. cerevisiae. The resulting strain had a 40% lower glycerol yield on glucose while the ethanol yield increased with 3% without affecting the maximum specific growth rate. Similarly, expression of GAPN in a strain harbouring xylose reductase and xylitol dehydrogenase led to an improvement in ethanol yield by up to 25% on xylose/glucose mixtures.

Computer Simulation↗

Integrating high-throughput and computational data elucidates bacterial networks.

The flood of high-throughput biological data has led to the expectation that computational (or in silico) models can be used to direct biological discovery, enabling biologists to reconcile heterogeneous data types, find inconsistencies and systematically generate hypotheses. Such a process is fundamentally iterative, where each iteration involves making model predictions, obtaining experimental data, reconciling the predicted outcomes with experimental ones, and using discrepancies to update the in silico model. Here we have reconstructed, on the basis of information derived from literature and databases, the first integrated genome-scale computational model of a transcriptional regulatory and metabolic network. The model accounts for 1,010 genes in Escherichia coli, including 104 regulatory genes whose products together with other stimuli regulate the expression of 479 of the 906 genes in the reconstructed metabolic network. This model is able not only to predict the outcomes of high-throughput growth phenotyping and gene expression experiments, but also to indicate knowledge gaps and identify previously unknown components and interactions in the regulatory and metabolic networks. We find that a systems biology approach that combines genome-scale experimentation and computation can systematically generate hypotheses on the basis of disparate data sources.

Aerobiosis↗

HUMESS: integrating quantitative transcriptomic analysis and metabolic modeling to unveil condition-specific gene signatures.

SUMMARY: Transcriptomic analysis is a key tool for exploring gene expression, but the complexity of biological systems often limits its insights. In particular, the lack of intermodal or multi-layered analysis hinders the ability to fully capture key cellular functions such as metabolism from transcriptomic data alone. Here, we introduce a novel approach that informs transcriptomic data analysis with metabolic network modeling to address this. Unlike traditional methods, HUman MEtabolism Specific Signature (HUMESS) uses genome-scale metabolic modeling and flux analysis to highlight reactions and involved genes based on their metabolic significance, offering a deeper understanding of transcriptomic data. Our computational pipeline, supported by a user-friendly Rshiny application, enhances gene expression analysis by uncovering metabolic phenotypic signatures. AVAILABILITY AND IMPLEMENTATION: HUMESS is open source and available under GitLab https://gitlab.univ-nantes.fr/bird_pipeline_registry/humess with the complete documentation available at https://gitlab.univ-nantes.fr/bird_pipeline_registry/humess/-/wikis/Home. A zenodo archive is also available at the following DOI: https://doi.org/10.5281/zenodo.15487717. An RShiny application has been developed to facilitate the exploration and analysis of HUMESS's results. The app is available online at the following address: https://shiny-bird.univ-nantes.fr/app/shinymess but can also be installed locally, available under GitLab https://gitlab.univ-nantes.fr/pare-l/shinymess.

Humans↗

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning↗

Constraints-based models: regulation of gene expression reduces the steady-state solution space.

Constraints-based models have been effectively used to analyse, interpret, and predict the function of reconstructed genome-scale metabolic models. The first generation of these models used "hard" non-adjustable constraints associated with network connectivity, irreversibility of metabolic reactions, and maximal flux capacities. These constraints restrict the allowable behaviors of a network to a convex mathematical solution space whose edges are extreme pathways that can be used to characterize the optimal performance of a network under a stated performance criterion. The development of a second generation of constraints-based models by incorporating constraints associated with regulation of gene expression was described in a companion paper published in this journal, using flux-balance analysis to generate time courses of growth and by-product secretion using a skeleton representation of core metabolism. The imposition of these additional restrictions prevents the use of a subset of the extreme pathways that are derived from the "hard" constraints, thus reducing the solution space and restricting allowable network functions. Here, we examine the reduction of the solution space due to regulatory constraints using extreme pathway analysis. The imposition of environmental conditions and regulatory mechanisms sharply reduces the number of active extreme pathways. This approach is demonstrated for the skeleton system mentioned above, which has 80 extreme pathways. As regulatory constraints are applied to the system, the number of feasible extreme pathways is reduced to between 26 and 2 extreme pathways, a reduction of between 67.5 and 97.5%. The method developed here provides a way to interpret how regulatory mechanisms are used to constrain network functions and produce a small range of physiologically meaningful behaviors from all allowable network functions.

Animals↗

In silico genome-scale reconstruction and validation of the Staphylococcus aureus metabolic network.

A genome-scale metabolic model of the Gram-positive, facultative anaerobic opportunistic pathogen Staphylococcus aureus N315 was constructed based on current genomic data, literature, and physiological information. The model comprises 774 metabolic processes representing approximately 23% of all protein-coding regions. The model was extensively validated against experimental observations and it correctly predicted main physiological properties of the wild-type strain, such as aerobic and anaerobic respiration and fermentation. Due to the frequent involvement of S. aureus in hospital-acquired bacterial infections combined with its increasing antibiotic resistance, we also investigated the clinically relevant phenotype of small colony variants and found that the model predictions agreed with recent findings of proteome analyses. This indicates that the model is useful in assisting future experiments to elucidate the interrelationship of bacterial metabolism and resistance. To help directing future studies for novel chemotherapeutic targets, we conducted a large-scale in silico gene deletion study that identified 158 essential intracellular reactions. A more detailed analysis showed that the biosynthesis of glycans and lipids is rather rigid with respect to circumventing gene deletions, which should make these areas particularly interesting for antibiotic development. The combination of this stoichiometric model with transcriptomic and proteomic data should allow a new quality in the analysis of clinically relevant organisms and a more rationalized system-level search for novel drug targets.

Computational Biology↗

A metabolic atlas of the Klebsiella pneumoniae species complex reveals lineage-specific metabolism and capacity for intra-species co-operation.

The Klebsiella pneumoniae species complex inhabits a wide variety of hosts and environments, and is a major cause of antimicrobial resistant infections. Genomics has revealed the population comprises multiple species/sub-species and hundreds of distinct co-circulating sub-lineage (SLs) that are associated with distinct gene complements. A substantial fraction of the pan-genome is predicted to be involved in metabolic functions and hence these data are consistent with metabolic differentiation at the SL level. However, this has so far remained unsubstantiated because in the past it was not possible to explore metabolic variation at scale. Here, we used a combination of comparative genomics and high-throughput genome-scale metabolic modeling to systematically explore metabolic diversity across the K. pneumoniae species complex (n = 7,835 genomes). We simulated growth outcomes for each isolate using carbon, nitrogen, phosphorus, and sulfur sources under aerobic and anaerobic conditions (n = 1,278 conditions per isolate). We showed that the distributions of metabolic genes and growth capabilities are structured in the population, and confirmed that SLs exhibit unique metabolic profiles. In vitro co-culture experiments demonstrated reciprocal commensalistic cross-feeding between SLs, effectively extending the range of conditions supporting individual growth. We propose that these substrate specializations may promote the existence and persistence of co-circulating SLs by reducing nutrient competition and facilitating commensal interactions. Our findings have implications for understanding the eco-evolutionary dynamics of K. pneumoniae and for the design of novel strategies to prevent opportunistic infections caused by this World Health Organization priority antimicrobial resistant pathogen.

Klebsiella pneumoniae↗

Metabolic network reconstruction as a resource for analyzing Salmonella Typhimurium SL1344 growth in the mouse intestine.

Nontyphoidal Salmonella strains (NTS) are among the most common foodborne enteropathogens and constitute a major cause of global morbidity and mortality, imposing a substantial burden on global health. The increasing antibiotic resistance of NTS bacteria has attracted a lot of research on understanding their modus operandi during infection. Growth in the gut lumen is a critical phase of the NTS infection. This might offer opportunities for intervention. However, the metabolic richness of the gut lumen environment and the inherent complexity and robustness of the metabolism of NTS bacteria call for modeling approaches to guide research efforts. In this study, we reconstructed a thermodynamically constrained and context-specific genome-scale metabolic model (GEM) for S. Typhimurium SL1344, a model strain well-studied in infection research. We combined sequence annotation, optimization methods and in vitro and in vivo experimental data. We used GEM to explore the nutritional requirements, the growth limiting metabolic genes, and the metabolic pathway usage of NTS bacteria in a rich environment simulating the murine gut. This work provides insight and hypotheses on the biochemical capabilities and requirements of SL1344 beyond the knowledge acquired through conventional sequence annotation and can inform future research aimed at better understanding NTS metabolism and identifying potential targets for infection prevention.

Salmonella typhimurium↗

An algorithmic framework for genome-wide modeling and analysis of translation networks.

The sequencing of genomes of several organisms and advances in high throughput technologies for transcriptome and proteome analysis has allowed detailed mechanistic studies of transcription and translation using mathematical frameworks that allow integration of both sequence-specific and kinetic properties of these fundamental cellular processes. To understand how perturbations in mRNA levels affect the synthesis of individual proteins within a large protein synthesis network, we consider here a genome-scale codon-wide model of the translation machinery with explicit description of the processes of initiation, elongation, and termination. The mechanistic codon-wide description of the translation process and the large number of mRNAs competing for resources, such as ribosomes, requires the use of novel efficient algorithmic approaches. We have developed such an efficient algorithmic framework for genome-scale models of protein synthesis. The mathematical and computational framework was applied to the analysis of the sensitivity of a translation network to perturbation in the rate constants and in the mRNA levels in the system. Our studies suggest that the highest specific protein synthesis rate (protein synthesis rate per mRNA molecule) is achieved when translation is elongation-limited. We find that the mRNA species with the highest number of actively translating ribosomes exerts maximum control on the synthesis of every protein, and the response of protein synthesis rates to mRNA expression variation is a function of the strength of initiation of translation at different mRNA species. Such quantitative understanding of the sensitivity of protein synthesis to the variation of mRNA expression can provide insights into cellular robustness mechanisms and guide the design of protein production systems.

Algorithms↗

Flux-sum coupling analysis of metabolic network models.

Metabolites acting as substrates and regulators of all biochemical reactions play an important role in maintaining the functionality of cellular metabolism. Despite advances in the constraint-based framework for genome-scale metabolic modeling, we lack reliable proxies for metabolite concentrations that can be efficiently determined and that allow us to investigate the relationship between metabolite concentrations in specific metabolic states in the absence of measurements. Here, we introduce a constraint-based approach, the flux-sum coupling analysis (FSCA), which facilitates the study of the interdependencies between metabolite concentrations by determining coupling relationships based on the flux-sum of metabolites. Application of FSCA on metabolic models of Escherichia coli, Saccharomyces cerevisiae, and Arabidopsis thaliana showed that the three coupling relationships are present in all models and pinpointed similarities in coupled metabolite pairs. Using the available concentration measurements of E. coli metabolites, we demonstrated that the coupling relationships identified by FSCA can capture the qualitative associations between metabolite concentrations and that flux-sum is a reliable proxy for metabolite concentration. Therefore, FSCA provides a novel tool for exploring and understanding the intricate interdependencies between the metabolite concentrations, advancing the understanding of metabolic regulation, and improving flux-centered systems biology approaches.

Escherichia coli↗

Toward whole cell modeling and simulation: comprehensive functional genomics through the constraint-based approach.

The increasing availability of various system-level, or so-called 'omics', datasets, in concert with existing data from the primary research literature, is facilitating the development of genome-scale metabolic models for many organisms. By incorporating the metabolic reaction stoichiometry as well as other physicochemical properties into systemic network reconstructions, these models account for the constraints that restrict an organism's phenotypic behavior. Accordingly, unlike many contemporary modeling strategies, this constraint-based modeling approach does not attempt to predict network behavior exactly; rather, it seeks to clearly distinguish those network states that a system can achieve from those that it cannot. A variety of analytical tools have been designed and developed to probe these models, thus enabling studies that investigate the metabolic capabilities of a number of organisms, that generate and test experimental hypotheses, and that predict accurately metabolic phenotypes and evolutionary outcomes. This chapter introduces the concepts that underlie the constraint-based modeling approach, and describes several of its applications with an emphasis on those potentially relevant to the drug development field. In addition, while this chapter focuses on the primary application of the constraint-based approach to date, namely in modeling metabolic networks, the latter sections of the chapter discuss its relatively recent application to modeling other cellular systems. Finally, the chapter concludes with an assessment of future directions focusing on the efforts that will be required to utilize the constraint-based approach in generating a holistic model of a viable organism.

Animals↗

High-Quality Genome Assembly, Metabolome, Pangenome, and Metabolic Models of Megasphaera hexanoica KCCM 43214T.

Megasphaera hexanoica KCCM 43214T, isolated from cow rumen, is capable of producing medium-chain carboxylic acids such as hexanoate and octanoate. In this study, we present a high-quality genome assembly, along with intracellular metabolomic profiling and pangenomic analysis. Illumina sequencing generated 2.3 Gbp from 15,293,634 reads with a GC content of 49.5%, while PacBio HiFi sequencing produced 331.5 Mbp across 45,266 reads, with an average read length of 7,323 bp and a HiFi read N50 of 8,214 bp. Hybrid assembly of short and long reads resulted in a single 2.88 Mbp contig, containing 2,835 protein-coding genes. Genome-scale metabolic models were constructed to evaluate its metabolic capabilities under specific growth conditions. Intracellular metabolomic analysis of cells grown in medium containing fructose and lactate revealed key metabolic activities associated with chain elongation. Pangenomic analysis across nine annotated genomes identified 6,721 orthologous genes using OrthoMCL, emphasizing the genetic and functional diversity within the Megasphaera genus. This dataset offers valuable insights into the metabolism and biotechnological potential of M. hexanoica KCCM 43214T.

Metabolome↗

Fine mapping of disease genes via haplotype clustering.

We propose an algorithm for analysing SNP-based population association studies, which is a development of that introduced by Molitor et al. [2003: Am J Hum Genet 73:1368-1384]. It uses clustering of haplotypes to overcome the major limitations of many current haplotype-based approaches. We define a between-haplotype score that is simple, yet appears to capture much of the information about evolutionary relatedness of the haplotypes in the vicinity of a (unobserved) putative causal locus. Haplotype clusters can then be defined via a putative ancestral haplotype and a cut-off distance. The number of an individual's two haplotypes that lie within the cluster predicts the individual's genotype at the causal locus. This predicted genotype can then be investigated for association with the phenotype of interest. We implement our approach within a Markov-chain Monte Carlo algorithm that, in effect, searches over locations and ancestral haplotypes to identify large, case-rich clusters. The algorithm successfully fine-maps a causal mutation in a test analysis using real data, and achieves almost 98% accuracy in predicting the genotype at the causal locus. A simulation study indicates that the new algorithm is substantially superior to alternative approaches, and it also allows us to identify situations in which multi-point approaches can substantially improve over single-SNP analyses. Our algorithm runs quickly and there is scope for extension to a wide range of disease models and genomic scales.

Algorithms↗

Metabolic functions of duplicate genes in Saccharomyces cerevisiae.

The roles of duplicate genes and their contribution to the phenomenon of enzyme dispensability are a central issue in molecular and genome evolution. A comprehensive classification of the mechanisms that may have led to their preservation, however, is currently lacking. In a systems biology approach, we classify here back-up, regulatory, and gene dosage functions for the 105 duplicate gene families of Saccharomyces cerevisiae metabolism. The key tool was the reconciled genome-scale metabolic model iLL672, which was based on the older iFF708. Computational predictions of all metabolic gene knockouts were validated with the experimentally determined phenotypes of the entire singleton yeast library of 4658 mutants under five environmental conditions. iLL672 correctly identified 96%-98% and 73%-80% of the viable and lethal singleton phenotypes, respectively. Functional roles for each duplicate family were identified by integrating the iLL672-predicted in silico duplicate knockout phenotypes, genome-scale carbon-flux distributions, singleton mutant phenotypes, and network topology analysis. The results provide no evidence for a particular dominant function that maintains duplicate genes in the genome. In particular, the back-up function is not favored by evolutionary selection because duplicates do not occur more frequently in essential reactions than singleton genes. Instead of a prevailing role, multigene-encoded enzymes cover different functions. Thus, at least for metabolism, persistence of the paralog fraction in the genome can be better explained with an array of different, often overlapping functional roles.

Gene Duplication↗

De Novo Genome Sequence Assembly of the Algal Endosymbiont Micractinium conductrix Derived From Its Host Paramecium bursaria 186b.

Endosymbiosis is a major driver of evolutionary innovation and underpins the function of diverse ecosystems. The origins and evolution of endosymbiosis are challenging to study experimentally due to the short-lived culturability of many microbial strains derived from endosymbiotic interactions. The facultative endosymbiosis between the ciliate, Paramecium bursaria, and the green alga, Micractinium conductrix (Chlorellaceae, Trebouxiophyceae), is ecologically widespread and has emerged as a powerful lab-tractable model system. This endosymbiosis is founded upon a reciprocal nutrient exchange, but each of the species can be cultured independently enabling quantification of symbiotic fitness effects, new partnerships to be generated in the lab, and co-associations to be subject to experimental evolution. To date, evolve-and-resequence approaches have been limited due to a lack of high-quality genome assemblies enabling gene variants to be identified. Here, we report a near telomere-to-telomere genome assembly for M. conductrix 186b, using a range of sequencing technologies. Comparative analysis shows that this is one of the most complete Chlorellaceae algal genome assemblies available to date. To aid accurate gene calling and annotation, we conducted both RNAseq and Iso-Seq transcriptome sequencing experiments. Collectively, these 'omics datasets will facilitate: (i) comparative genomics studies of endosymbiont evolution, (ii) evolve-and-resequence experiments, (iii) genome-scale metabolic modeling studies, and (iv) identification of targets for genetic modification experiments and biotechnological applications.

Symbiosis↗