Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Conservation of expression and sequence of metabolic genes is reflected by activity across metabolic states.

Variation in gene expression levels on a genomic scale has been detected among different strains, among closely related species, and within populations of genetically identical cells. What are the driving forces that lead to expression divergence in some genes and conserved expression in others? Here we employ flux balance analysis to address this question for metabolic genes. We consider the genome-scale metabolic model of Saccharomyces cerevisiae, and its entire space of optimal and near-optimal flux distributions. We show that this space reveals underlying evolutionary constraints on expression regulation, as well as on the conservation of the underlying gene sequences. Genes that have a high range of optimal flux levels tend to display divergent expression levels among different yeast strains and species. This suggests that gene regulation has diverged in those parts of the metabolic network that are less constrained. In addition, we show that genes that are active in a large fraction of the space of optimal solutions tend to have conserved sequences. This supports the possibility that there is less selective pressure to maintain genes that are relevant for only a small number of metabolic states.

Cell Proliferation↗

Understanding disease-associated metabolic changes in human colonic epithelial cells using the iColonEpithelium metabolic reconstruction.

The colonic epithelium plays a key role in the host-microbiome interactions, allowing uptake of various nutrients and driving important metabolic processes. To unravel detailed metabolic activities in the human colonic epithelium, our present study focuses on the generation of the first cell-type-specific genome-scale metabolic model (GEM) of human colonic epithelial cells, named iColonEpithelium. GEMs are powerful tools for exploring reactions and metabolites at the systems level and predicting the flux distributions at steady state. Our cell-type-specific iColonEpithelium metabolic reconstruction captures genes specifically expressed in the human colonic epithelial cells. iColonEpithelium is also capable of performing metabolic tasks specific to the colonic epithelium. A unique transport reaction compartment has been included to allow for the simulation of metabolic interactions with the gut microbiome. We used iColonEpithelium to identify metabolic signatures associated with inflammatory bowel disease. We used single-cell RNA sequencing data from Crohn's Diseases (CD) and ulcerative colitis (UC) samples to build disease-specific iColonEpithelium metabolic networks in order to predict metabolic signatures of colonocytes in both healthy and disease states. We identified reactions in nucleotide interconversion, fatty acid synthesis and tryptophan metabolism were differentially regulated in CD and UC conditions, relative to healthy control, which were in accordance with experimental results. The iColonEpithelium metabolic network can be used to identify mechanisms at the cellular level, and we show an initial proof-of-concept for how our tool can be leveraged to explore the metabolic interactions between host and gut microbiota.

Humans↗

Computational modeling of the Plasmodium falciparum interactome reveals protein function on a genome-wide scale.

Many thousands of proteins encoded by the genome of Plasmodium falciparum, the causal organism of the deadliest form of human malaria, are of unknown function. It is of utmost importance that these proteins be characterized if we are to develop combative strategies against malaria based on the biology of the parasite. In an attempt to infer protein function on a genome-wide scale, we computationally modeled the P. falciparum interactome, elucidating local and global functional relationships between gene products. The resulting interaction network, reconstructed by integrating in silico and experimental functional genomics data within a Bayesian framework, covers approximately 68% of the parasite genome and provides functional inferences for more than 2000 uncharacterized proteins, based on their associations. Network reconstruction involved the use of a novel strategy, where we incorporated continuously updated, uniform reference priors in our Bayesian model. This method for generating interaction maps is thus also well suited for application to other genomes, where pre-existing interactome knowledge is sparse. Additionally, we superimposed this map on genomes of three apicomplexan pathogens--Plasmodium yoelii, Toxoplasma gondii, and Cryptosporidium parvum--describing relationships between these organisms based on retained functional linkages. This comparison provided a glimpse of the highly evolved nature of P. falciparum; for instance, a deficit of nearly 26% in terms of predicted interactions is observed against P. yoelii, because of missing ortholog partners in pairs of functionally linked proteins.

Animals↗

Large-scale evaluation of in silico gene deletions in Saccharomyces cerevisiae.

A large-scale in silico evaluation of gene deletions in Saccharomyces cerevisiae was conducted using a genome-scale reconstructed metabolic model. The effect of 599 single gene deletions on cell viability was simulated in silico and compared to published experimental results. In 526 cases (87.8%), the in silico results were in agreement with experimental observations when growth on synthetic complete medium was simulated. Viable phenotypes were correctly predicted in 89.4% (496 out of 555) and lethal phenotypes were correctly predicted in 68.2% (30 out of 44) of the cases considered. The in silico evaluation was solely based on the topological properties of the metabolic network which is based on well-established reaction stoichiometry. No interaction or regulatory information was accounted for in the in silico model. False predictions were analyzed on a case-by-case basis for four possible inadequacies of the in silico model: (1) incomplete media composition, (2) substitutable biomass components, (3) incomplete biochemical information, and (4) missing regulation. This analysis eliminated a number of false predictions and suggested a number of experimentally testable hypotheses. A genome-scale in silico model can thus be used to systematically reconcile existing data and fill in our knowledge gaps about an organism.

Computational Biology↗

Systematic analysis of conservation relations in Escherichia coli genome-scale metabolic network reveals novel growth media.

A biochemical species is called producible in a constraints-based metabolic model if a feasible steady-state flux configuration exists that sustains its nonzero concentration during growth. Extreme semipositive conservation relations (ESCRs) are the simplest semipositive linear combinations of species concentrations that are invariant to all metabolic flux configurations. In this article, we outline a fundamental relationship between the ESCRs of a metabolic network and the producibility of a biochemical species under a nutrient media. We exploit this relationship in an algorithm that systematically enumerates all minimal nutrient sets that render an objective species weakly producible (i.e., producible in the absence of thermodynamic constraints) through a simple traversal of ESCRs. We apply our results to a recent genome scale model of Escherichia coli metabolism, in which we traverse the 51 anhydrous ESCRs of the metabolic network to determine all 928 minimal aqueous nutrient media that render biomass weakly producible. Applying irreversibility constraints, we find 287 of these 928 nutrient sets to be thermodynamically feasible. We also find that an additional 365 of these nutrient sets are thermodynamically feasible in the presence of oxygen. Since biomass producibility is commonly used as a surrogate for growth in genome scale metabolic models, our results represent testable hypotheses of alternate growth media derived from in silico analysis of the E. coli genome scale metabolic network.

Algorithms↗

SimHumanity: Using SLiM 5.0 to run whole-genome simulations of human evolution.

The reconstruction of human evolutionary history has undergone repeated advances, each made possible by methodological innovations. In recent decades, genetic and genomic data played a central role in the reconstruction of major evolutionary events such as the out-of-Africa migration, and genetic simulations of human evolutionary history have come to play a major role in testing more specific hypotheses including proposed patterns of migration and admixture with archaic hominins. Increasing computational power has allowed human evolutionary history to be modeled at ever-larger scales, but simulations that encompass the complete human genome, including sex chromosomes and mitochondrial DNA, have been difficult due to the lack of support for whole-genome models in commonly used evolutionary simulation frameworks. With the recent introduction of SLiM 5 such simulations are now straightforward to construct, allowing the easy simulation of humans at whole-genome scale under different demographic models and evolutionary dynamics. We here present three versions of a reusable, customizable, open-source SLiM 5 model for simulating the molecular evolution of the full human genome. We also show some simple analyses of results from the model, to illustrate its utility. We hope this model, which we have nicknamed "SimHumanity" in jest, will facilitate further progress in the field of human evolutionary simulations.

SLiM↗

Genome-scale reconstruction of the metabolic network in Staphylococcus aureus N315: an initial draft to the two-dimensional annotation.

BACKGROUND: Several strains of bacteria have sequenced and annotated genomes, which have been used in conjunction with biochemical and physiological data to reconstruct genome-scale metabolic networks. Such reconstruction amounts to a two-dimensional annotation of the genome. These networks have been analyzed with a constraint-based formalism and a variety of biologically meaningful results have emerged. Staphylococcus aureus is a pathogenic bacterium that has evolved resistance to many antibiotics, representing a significant health care concern. We present the first manually curated elementally and charge balanced genome-scale reconstruction and model of S. aureus' metabolic networks and compute some of its properties. RESULTS: We reconstructed a genome-scale metabolic network of S. aureus strain N315. This reconstruction, termed iSB619, consists of 619 genes that catalyze 640 metabolic reactions. For 91% of the reactions, open reading frames are explicitly linked to proteins and to the reaction. All but three of the metabolic reactions are both charge and elementally balanced. The reaction list is the most complete to date for this pathogen. When the capabilities of the reconstructed network were analyzed in the context of maximal growth, we formed hypotheses regarding growth requirements, the efficiency of growth on different carbon sources, and potential drug targets. These hypotheses can be tested experimentally and the data gathered can be used to improve subsequent versions of the reconstruction. CONCLUSION: iSB619 represents comprehensive biochemically and genetically structured information about the metabolism of S. aureus to date. The reconstructed metabolic network can be used to predict cellular phenotypes and thus advance our understanding of a troublesome pathogen.

Bacterial Proteins↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

Genomic approaches in dissecting complex biological pathways.

Advances in genomic research have provided many types of large-scale data that contain rich information on various biological pathways. Intensive efforts have been made to qualitatively or quantitatively model biological pathways using these genomic data. Some general network properties, such as the scale-free property and network motifs, have been discussed and various network models have been applied to reconstruct pathways. However, there is a lack of systematic integration of prior knowledge and different genomic data in these analyses. In this review, we discuss pathway reconstruction under the consideration of the complexity embedded in the biological system, and the global and local properties of biological pathways. We review major methodologies, including clustering methods, scale-free networks models, Bayesian networks models, Boolean networks models, systems of differential equations, and data integration methods. We focus on the difficulty of each methodology in modeling biological pathways, and emphasize that different models capture different aspects of biological pathways or genomic data. The 'noisy' large-scale genomic data require the mathematical models and computational methods to be both robust and identifiable. In addition, we believe that ideal models should have the capability of incorporating various data types and these models need to be assessed through rigorous comparisons with empirical data.

Bayes Theorem↗

Genome-scale analysis of the uses of the Escherichia coli genome: model-driven analysis of heterogeneous data sets.

The recent availability of heterogeneous high-throughput data types has increased the need for scalable in silico methods with which to integrate data related to the processes of regulation, protein synthesis, and metabolism. A sequence-based framework for modeling transcription and translation in prokaryotes has been established and has been extended to study the expression state of the entire Escherichia coli genome. The resulting in silico analysis of the expression state highlighted three facets of gene expression in E. coli: (i) the metabolic resources required for genome expression and protein synthesis were found to be relatively invariant under the conditions tested; (ii) effective promoter strengths were estimated at the genome scale by using global mRNA abundance and half-life data, revealing genes subject to regulation under the experimental conditions tested; and (iii) large-scale genome location-dependent expression patterns with approximately 600-kb periodicity were detected in the E. coli genome based on the 49 expression data sets analyzed. These results support the notion that a structured model-driven analysis of expression data yields additional information that can be subjected to commonly used statistical analyses. The integration of heterogeneous genome-scale data (i.e., sequence, expression data, and mRNA half-life data) is readily achieved in the context of an in silico model.

Bacterial Proteins↗

Comparative protein structure modeling of genes and genomes.

Comparative modeling predicts the three-dimensional structure of a given protein sequence (target) based primarily on its alignment to one or more proteins of known structure (templates). The prediction process consists of fold assignment, target-template alignment, model building, and model evaluation. The number of protein sequences that can be modeled and the accuracy of the predictions are increasing steadily because of the growth in the number of known protein structures and because of the improvements in the modeling software. Further advances are necessary in recognizing weak sequence-structure similarities, aligning sequences with structures, modeling of rigid body shifts, distortions, loops and side chains, as well as detecting errors in a model. Despite these problems, it is currently possible to model with useful accuracy significant parts of approximately one third of all known protein sequences. The use of individual comparative models in biology is already rewarding and increasingly widespread. A major new challenge for comparative modeling is the integration of it with the torrents of data from genome sequencing projects as well as from functional and structural genomics. In particular, there is a need to develop an automated, rapid, robust, sensitive, and accurate comparative modeling pipeline applicable to whole genomes. Such large-scale modeling is likely to encourage new kinds of applications for the many resulting models, based on their large number and completeness at the level of the family, organism, or functional network.

Animals↗

Protein fragment reconstruction using various modeling techniques.

Recently developed reduced models of proteins with knowledge-based force fields have been applied to a specific case of comparative modeling. From twenty high resolution protein structures of various structural classes, significant fragments of their chains have been removed and treated as unknown. The remaining portions of the structures were treated as fixed - i.e., as templates with an exact alignment. Then, the missed fragments were reconstructed using several modeling tools. These included three reduced types of protein models: the lattice SICHO (Side Chain Only) model, the lattice CABS (Calpha + Cbeta + Side group) model and an off-lattice model similar to the CABS model and called REFINER. The obtained reduced models were compared with more standard comparative modeling tools such as MODELLER and the SWISS-MODEL server. The reduced model results are qualitatively better for the higher resolution lattice models, clearly suggesting that these are now mature, competitive and complementary (in the range of sparse alignments) to the classical tools of comparative modeling. Comparison between the various reduced models strongly suggests that the essential ingredient for the sucessful and accurate modeling of protein structures is not the representation of conformational space (lattice, off-lattice, all-atom) but, rather, the specificity of the force fields used and, perhaps, the sampling techniques employed. These conclusions are encouraging for the future application of the fast reduced models in comparative modeling on a genomic scale.

Amino Acid Sequence↗

Properties of metabolic networks: structure versus function.

Biological data from high-throughput technologies describing the network components (genes, proteins, metabolites) and their associated interactions have driven the reconstruction and study of structural (topological) properties of large-scale biological networks. In this article, we address the relation of the functional and structural properties by using extensively experimentally validated genome-scale metabolic network models to compute observable functional states of a microorganism and compare the "structure versus function" attributes of metabolic networks. It is observed that, functionally speaking, the essentiality of reactions in a node is not correlated with node connectivity as structural analyses of other biological networks have suggested. These findings are illustrated with the analysis of the genome-scale biochemical networks of three species with distinct modes of metabolism. These results also suggest fundamental differences among different biological networks arising out of their representation and functional constraints.

Algorithms↗

Integrated analysis of regulatory and metabolic networks reveals novel regulatory mechanisms in Saccharomyces cerevisiae.

We describe the use of model-driven analysis of multiple data types relevant to transcriptional regulation of metabolism to discover novel regulatory mechanisms in Saccharomyces cerevisiae. We have reconstructed the nutrient-controlled transcriptional regulatory network controlling metabolism in S. cerevisiae consisting of 55 transcription factors regulating 750 metabolic genes, based on information in the primary literature. This reconstructed regulatory network coupled with an existing genome-scale metabolic network model allows in silico prediction of growth phenotypes of regulatory gene deletions as well as gene expression profiles. We compared model predictions of gene expression changes in response to genetic and environmental perturbations to experimental data to identify potential novel targets for transcription factors. We then identified regulatory cascades connecting transcription factors to the potential targets through a systematic model expansion strategy using published genome-wide chromatin immunoprecipitation and binding-site-motif data sets. Finally, we show the ability of an integrated metabolic and regulatory network model to predict growth phenotypes of transcription factor knockout strains. These studies illustrate the potential of model-driven data integration to systematically discover novel components and interactions in regulatory and metabolic networks in eukaryotic cells.

Chromatin Immunoprecipitation↗

Comparative analysis of plant genome architecture.

Many genes are similar in most plants and it is clear that the ordering of genes is highly conserved across wide taxonomic groupings. Repetitive DNA, consisting of sequence motifs between 2 and 10,000 base pairs long, repeated many hundreds or thousands of times in the genome, represents the majority of most plant genomes and defines some of the differences between species. Some sequences are highly conserved in many species, while other sequences show species or even chromosome specificity. Different types of sequences have markedly contrasting genomic distributions; even among tandem repeats, some are sub-terminal, some paracentromeric and others intercalary. The reasons for these different distributions are largely unknown, and mechanisms of homogenization, dispersion and amplifications are the subject of much speculation. Aspects of plant genome architecture-the organization of repetitive and single-copy DNA sequences along the chromosomes, and the positioning of those sequences within the nucleus at interphase-have important consequences for plant genetics. Models of large scale genome organization may be useful in learning the function of different components of the genome, in evolutionary studies and in plant breeding.

Biological Evolution↗

High-throughput mutation detection underlying adaptive evolution of Escherichia coli-K12.

Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) analysis of base-specific cleavage products is an efficient, highly accurate tool for the detection of single base sequence variations. We describe the first application of this comparative sequencing strategy for automated high-throughput mutation detection in microbial genomes. The method was applied to identify DNA sequence changes that occurred in Escherichia coli K-12 MG1655 during laboratory adaptive evolution to new optimal growth phenotypes. Experiments were based on a genome-scale in silico model of E. coli metabolism and growth. This model computes several phenotypic functions and predicts optimal growth rates. To identify mutations underlying a 40-d adaptive laboratory evolution on glycerol, we resequenced 4.4% of the E. coli-K12 MG1655 genome in several clones picked at the end of the evolutionary process. The 1.54-Mb screen was completed in 13.5 h. This resequencing study is the largest reported by MALDI-TOF mass spectrometry to date. Ten mutations in 40 clones and three deviations from the reference sequence were detected. Mutations were predominantly found within the glycerol kinase gene. Functional characterization of the most prominent mutation shows its metabolic impact on the process of adaptive evolution. All sequence changes were independently confirmed by genotyping and Sanger-sequencing. We demonstrate that comparative sequencing by base-specific cleavage and MALDI-TOF mass spectrometry is an automated, fast, and highly accurate alternative to capillary sequencing.

Adaptation, Biological↗

Large-scale protein structure modeling of the Saccharomyces cerevisiae genome.

The function of a protein generally is determined by its three-dimensional (3D) structure. Thus, it would be useful to know the 3D structure of the thousands of protein sequences that are emerging from the many genome projects. To this end, fold assignment, comparative protein structure modeling, and model evaluation were automated completely. As an illustration, the method was applied to the proteins in the Saccharomyces cerevisiae (baker's yeast) genome. It resulted in all-atom 3D models for substantial segments of 1,071 (17%) of the yeast proteins, only 40 of which have had their 3D structure determined experimentally. Of the 1,071 modeled yeast proteins, 236 were related clearly to a protein of known structure for the first time; 41 of these previously have not been characterized at all.

Fungal Proteins↗

High-throughput fluorescence microscopy for systems biology.

In this post-genomic era, we need to define gene function on a genome-wide scale for model organisms and humans. The fundamental unit of biological processes is the cell. Among the most powerful tools to assay such processes in the physiological context of intact living cells are fluorescence microscopy and related imaging techniques. To enable these techniques to be applied to functional genomics experiments, fluorescence microscopy is making the transition to a quantitative and high-throughput technology.

Animals↗