Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Predicting metal-binding site residues in low-resolution structural models.

The accurate prediction of the biochemical function of a protein is becoming increasingly important, given the unprecedented growth of both structural and sequence databanks. Consequently, computational methods are required to analyse such data in an automated manner to ensure genomes are annotated accurately. Protein structure prediction methods, for example, are capable of generating approximate structural models on a genome-wide scale. However, the detection of functionally important regions in such crude models, as well as structural genomics targets, remains an extremely important problem. The method described in the current study, MetSite, represents a fully automatic approach for the detection of metal-binding residue clusters applicable to protein models of moderate quality. The method involves using sequence profile information in combination with approximate structural data. Several neural network classifiers are shown to be able to distinguish metal sites from non-sites with a mean accuracy of 94.5%. The method was demonstrated to identify metal-binding sites correctly in LiveBench targets where no obvious metal-binding sequence motifs were detectable using InterPro. Accurate detection of metal sites was shown to be feasible for low-resolution predicted structures generated using mGenTHREADER where no side-chain information was available. High-scoring predictions were observed for a recently solved hypothetical protein from Haemophilus influenzae, indicating a putative metal-binding site.

Binding Sites↗

Modelling in molecular biology: describing transcription regulatory networks at different scales.

Approaches to describe gene regulation networks can be categorized by increasing detail, as network parts lists, network topology models, network control logic models or dynamic models. We discuss the current state of the art for each of these approaches. We study the relationship between different topology models, and give examples how they can be used to infer functional annotations for genes of unknown function. We introduce a new simple way of describing dynamic models called finite state linear model (FSLM). We discuss the gap between the parts list and topology models on one hand, and network logic and dynamic models, on the other hand. The first two classes of models have reached a genome-wide scale, while for the other model classes high-throughput technologies are yet to make a major impact.

Computational Biology↗

Genome-scale in silico aided metabolic analysis and flux comparisons of Escherichia coli to improve succinate production.

In the post-genome era, it is one challenge to understand the cellular metabolism at the systematic levels. Mathematical modeling of microorganisms and subsequent computer simulation are effective tools for systems biology. In this paper, based on the genome-scale Escherichia coli stoichiometric model iJR904, through the GAMS linear programming package, the in silico maximal succinate yield was estimated to be 1.714 mol/mol glucose. When another two constraints were added, the maximal succinate yield dropped to 1.60 mol/mol glucose. Further analysis substantiated the uniqueness of the flux distribution under such constraints. After comparisons with the metabolic flux analysis (MFA) results computed from the wet experimental data of the three kinds of E. coli, three potential improvement target sites, the glucose phosphotransferase transport system, the pyruvate carboxylase, and the glyoxylate shunt, were identified and selected for the genetic modifications. All the three genetic modified strains showed increased succinate yield. The final strain TUQ19/pQZ6 had a high yield of 1.29 mol succinate/mol glucose and high productivity. The success of the above experiments proved that this in silico optimal succinate production pathway is reasonable and practical. This method may also be used as a general strategy to help enhance the yields of other favorable metabolites in E. coli.

Anaerobiosis↗

Modelling gene networks at different organisational levels.

Approaches to modelling gene regulation networks can be categorized, according to increasing detail, as network parts lists, network topology models, network control logic models, or dynamic models. We discuss the current state of the art for each of these approaches. There is a gap between the parts list and topology models on one hand, and control logic and dynamic models on the other hand. The first two classes of models have reached a genome-wide scale, while for the other model classes high throughput technologies are yet to make a major impact.

Algorithms↗

Selection favors cost-ordered adaptive mutational paths in a stress-magnitude-dependent manner.

Chronic exposure to stress requires adaptive strategies beyond canonical regulatory mechanisms. Stress varies, both qualitatively and quantitatively, across physiological niches and exerts distinct selection pressures on colonizing bacteria. Bacteria employ diverse defense strategies to withstand various stressful conditions, yet how they tailor their responses to different magnitudes of the same stressor remains poorly understood. We used multiple adaptive laboratory evolution experiments of Escherichia coli across varying paraquat concentrations and genetic backgrounds to dissect adaptive strategies at different levels of stress. Integrating multi-omic analyses with a tailored genome-scale metabolic model-based parametrized cost calculations, we identify two fundamentally distinct tolerance mechanisms. Under low-paraquat stress, blocking the polyamine transporter that is reported to be hijacked for paraquat influx suffices to maintain optimal growth. In contrast, higher stress levels activate an energetically demanding program involving enhanced detoxification and efflux. The transport flux regulation establishes a primary defense layer, upon which metabolic repair systems provide additional fitness advantages. The stress magnitude-dependent differential engagement of previously reported paraquat tolerance approaches offers insights into the principles governing dynamic bacterial adaptation.

Journal Article↗

Targeted, Genome-scale Overexpression in Proteobacteria.

Targeted, genome-scale gene perturbation screens using Clustered Regularly Interspaced Short Palindromic Repeats interference (CRISPRi) and activation (CRISPRa) have revolutionized eukaryotic genetics, advancing medical, industrial, and basic research. Although CRISPRi knockdowns have been broadly applied in bacteria, options for genome-scale gene overexpression face key limitations. Here, we develop a facile approach for genome-scale overexpression in bacteria we call, "CRISPRtOE" (CRISPR transposition and OverExpression). We first create a platform for comprehensive gene targeting using CRISPR-associated transposons (CAST) and show that transposition occurs at a higher frequency in non-transcribed DNA. We then demonstrate that CRISPRtOE can upregulate gene expression in Proteobacteria with medical and industrial relevance by integrating synthetic promoters of varying strength upstream of target genes. Finally, we employ CRISPRtOE screening at the genome-scale in the model bacterium Escherichia coli and the non-model biofuel producer Zymomonas mobilis, recovering known and novel antibiotic and engineering targets. We envision that CRISPRtOE will be a valuable overexpression tool for antibiotic mode of action, industrial strain optimization, and gene function discovery in bacteria.

Journal Article↗

Perspectives: sequence data base searching in the era of large-scale genomic sequencing.

Large-scale sequencing of human and model organism genomes will have a profound impact on our ability to use sequence data base searching to predict the biochemical functions of sequences of interest. Despite the great value of more sequences in the data bases, a huge increase in data base size will also have adverse effects on data base searches. Upcoming problems will include (1) greatly increased search times, (2) an increase in background noise of high-scoring but biologically irrelevant matches, (3) inaccurate coding region prediction, leading to problems in protein data base searching, and (4) limited first-pass sequence annotation, making it difficult to determine the biological relevance of data base hits. Improved data base annotation tools and construction of smaller data bases of representative and highly-annotated sequences for first-pass analyses will be essential to deal with the impending flood of new genomic sequence.

Animals↗

Regulation of gene expression in flux balance models of metabolism.

Genome-scale metabolic networks can now be reconstructed based on annotated genomic data augmented with biochemical and physiological information about the organism. Mathematical analysis can be performed to assess the capabilities of these reconstructed networks. The constraints-based framework, with flux balance analysis (FBA), has been used successfully to predict time course of growth and by-product secretion, effects of mutation and knock-outs, and gene expression profiles. However, FBA leads to incorrect predictions in situations where regulatory effects are a dominant influence on the behavior of the organism. Thus, there is a need to include regulatory events within FBA to broaden its scope and predictive capabilities. Here we represent transcriptional regulatory events as time-dependent constraints on the capabilities of a reconstructed metabolic network to further constrain the space of possible network functions. Using a simplified metabolic/regulatory network, growth is simulated under various conditions to illustrate systemic effects such as catabolite repression, the aerobic/anaerobic diauxic shift and amino acid biosynthesis pathway repression. The incorporation of transcriptional regulatory events in FBA enables us to interpret, analyse and predict the effects of transcriptional regulation on cellular metabolism at the systemic level.

Amino Acids↗

Mitochondrial genome expression in a mutant strain of D. subobscura, an animal model for large scale mtDNA deletion.

A mitochondrial mutant strain of D. subobscura has two mitochondrial genome populations (heteroplasmy): the first (20-30% of the population, 15.9 kb) is the same as could be found in the wild type; the second (70-80% of the population, 11 kb) has lost by deletion several genes coding for complex I and III subunits, and four tRNAs. In human pathology, this kind of mutation has been correlated with severe diseases such as the Kearns-Sayre syndrome, but the mutant strain, does not seem to be affected by the mutation (1). Studies reported here show that: a) Transcripts from genes not concerned by the mutation are present at the same level in both strains. b) In contrast, transcript concentrations from genes involved in the deletion are significantly decreased (30-50%) in the mutant. c) Deleted DNA was expressed as shown by the detection of the fusion transcript. d) The mtDNA/nuc.DNA ratio is 1.5 times higher in the mutant strain than in the wild type. The mutation leads to change in the transcript level equilibrium. The apparent innocuousness of the mutation may suggest some post-transcriptional compensation mechanisms. This drosophila strain is an interesting model to study the consequence of this type of mitochondrial genome deletion.

Animals↗

Perspectives for vascular genomics.

Diseases of the vascular system result from a complex mixture of genetic and environmental factors. Data sets, technologies and strategies emanating from the human genome programme have been applied to the analysis of both rare single-gene and common multigenic vascular disorders. Genomic approaches including inter- and intraspecies sequence comparisons, genotyping with dense marker sets spanning the genome, large-scale mutagenesis screens of model organisms, and genome-wide expression profiling have all begun to contribute to the identification of new genes and mechanisms that are central to cardiovascular disease processes.

Animals↗

ReMeDy: A Flexible Statistical Framework for Region-Based Detection of DNA Methylation Dysregulation.

Region-based epigenome-wide association studies have demonstrated improved statistical power and biological interpretability compared with probe-wise analyses of DNA methylation data. However, most existing region-based methods characterize methylation dysregulation primarily through changes in mean methylation levels associated with a phenotype of interest. Substantial evidence indicates that phenotype-associated methylation alterations may also manifest through changes in methylation variability or through joint shifts in mean and variability. Despite this, no existing statistical framework jointly models mean-variance methylation changes in a region-based manner. We propose ReMeDy, a flexible statistical framework that uses a hierarchical likelihood approach within a generalized linear model setting to identify differentially methylated regions, variably methylated regions, and regions exhibiting joint differential and variable methylation at a genome-wide scale. Unlike existing models, ReMeDy operates directly on biologically defined co-methylated regions, allowing it to naturally capture spatial correlation inherent in DNA methylation array data, while avoiding reliance on heuristic, user-defined tuning parameters such as smoothing spans and kernel bandwidths that can substantially influence results and introduce subjectivity. Through extensive simulation studies and comprehensive benchmarking against popular models, we demonstrate that ReMeDy maintains false discovery and Type-I error rates at nominal levels while achieving consistently higher statistical power across a wide range of realistic scenarios. Application to population-level DNA methylation data further shows that ReMeDy identifies biologically meaningful regions and pathways implicated in complex human diseases that are not captured by conventional mean-based analyses alone. ReMeDy is implemented as an open-source R package and is freely available at https://github.com/SChatLab/ReMeDy.

DNA Methylation↗

Prediction of nuclear hormone receptor response elements.

The nuclear receptor (NR) class of transcription factors controls critical regulatory events in key developmental processes, homeostasis maintenance, and medically important diseases and conditions. Identification of the members of a regulon controlled by a NR could provide an accelerated understanding of development and disease. New bioinformatics methods for the analysis of regulatory sequences are required to address the complex properties associated with known regulatory elements targeted by the receptors because the standard methods for binding site prediction fail to reflect the diverse target site configurations. We have constructed a flexible Hidden Markov Model framework capable of predicting NHR binding sites. The model allows for variable spacing and orientation of half-sites. In a genome-scale analysis enabled by the model, we show that NRs in Fugu rubripes have a significant cross-regulatory potential. The model is implemented in a web interface, freely available for academic researchers, available at http://mordor.cgb.ki.se/NHR-scan.

Algorithms↗

Network topology and the evolution of dynamics in an artificial genetic regulatory network model created by whole genome duplication and divergence.

Topological measures of large-scale complex networks are applied to a specific artificial regulatory network model created through a whole genome duplication and divergence mechanism. This class of networks share topological features with natural transcriptional regulatory networks. Specifically, these networks display scale-free and small-world topology and possess subgraph distributions similar to those of natural networks. Thus, the topologies inherent in natural networks may be in part due to their method of creation rather than being exclusively shaped by subsequent evolution under selection. The evolvability of the dynamics of these networks is also examined by evolving networks in simulation to obtain three simple types of output dynamics. The networks obtained from this process show a wide variety of topologies and numbers of genes indicating that it is relatively easy to evolve these classes of dynamics in this model.

Computational Biology↗

Metabolism and gene expression models for the microbiome reveal how diet and metabolic dysbiosis impact disease.

The gut microbiome plays a critical role in human health, spurring extensive research using multi-omic technologies. Although these tools offer valuable insights, they often fall short in capturing the complexity of microbial interactions that associate with disease onset, progression, and treatment. Thus, integration of multi-omics datasets with metabolic models is needed to predict associations between microbial activity and disease. Here, we automated the reconstruction of 495 metabolic and gene expression models (ME-models), overcoming the main limitation preventing the wide use of this approach. We integrated them with multi-omics data from patients with inflammatory bowel disease (IBD), identifying taxa associated with variations in amino acids, short-chain fatty acids, and pH in the gut of IBD patients. In general, this approach provides testable hypotheses of the metabolic activity of the gut microbiota, and the automated pipeline opens the opportunity to study microbial interactions in other biologically relevant settings using ME-models.

Humans↗

Biophysical metabolic modeling of complex bacterial colony morphology.

Microbial colony growth is shaped by the physics of biomass propagation and nutrient diffusion and by the metabolic reactions that organisms activate as a function of the surrounding environment. While microbial colonies have been explored using minimal models of growth and motility, full integration of biomass propagation and metabolism is still lacking. Here, building upon our framework for computation of microbial ecosystems in time and space (COMETS), we combine dynamic flux balance modeling of metabolism with collective biomass propagation and demographic fluctuations to provide nuanced simulations of E. coli colonies. Simulations produced realistic colony morphology, consistent with our experiments. They characterize the transition between smooth and furcated colonies and the decay of genetic diversity. Furthermore, we demonstrate that under certain conditions, biomass can accumulate along "metabolic rings" that are reminiscent of coffee-stain rings but have a completely different origin. Our approach is a key step toward predictive microbial ecosystems modeling. A record of this paper's transparent peer review process is included in the supplemental information.

Models, Biological↗

Analysis of B cell memory formation using DNA microarrays.

DNA microarray analysis of B cell subsets has identified comprehensive programs of gene expression that distinguish B cells at discrete stages of differentiation. The next task is to identify key genetic signals within these complex programs that regulate the dynamic cellular events during B cell activation in vivo. After stimulation with antigen, naïve B cells proliferate and differentiate, and then produce antibodies. Crucial qualitative differences in antibody responses are observed depending on whether or not B cells receive T cell help during activation. Proteins, lipopolysaccharides, and polysaccharides stimulate T-dependent (TD), T-independent type 1 (TI-1), and type 2 (TI-2) antibody responses, respectively. Only TD responses generate somatically mutated antibody-forming (plasma) cells and memory B cells, which produce high affinity anamnestic responses to subsequent antigen challenge. Somatic mutation of immunoglobulin genes occurs during B cell proliferation in germinal centres (GC), which are typical in TD responses but rare in TI responses. However, we have described a model, which is exceptional because numerous large GC form in response to a model TI-2 antigen, (4-hydoxy-3-nitrophenyl) acetyl (NP)-Ficoll. Significantly, these GC undergo involution before memory B cells are generated. This model provides an opportunity to investigate the genetic signals that drive memory cell formation, and we have compared global gene expression in TI and TD GC to identify a relatively small number of genes that are differentially expressed between the two prototypic B cell responses. This model demonstrates how genome-scale technology can be adapted to investigate specific aspects of B cell biology.

Animals↗

The global transcriptional regulatory network for metabolism in Escherichia coli exhibits few dominant functional states.

A principal aim of systems biology is to develop in silico models of whole cells or cellular processes that explain and predict observable cellular phenotypes. Here, we use a model of a genome-scale reconstruction of the integrated metabolic and transcriptional regulatory networks for Escherichia coli, composed of 1,010 gene products, to assess the properties of all functional states computed in 15,580 different growth environments. The set of all functional states of the integrated network exhibits a discernable structure that can be visualized in 3-dimensional space, showing that the transcriptional regulatory network governing metabolism in E. coli responds primarily to the available electron acceptor and the presence of glucose as the carbon source. This result is consistent with recently published experimental data. The observation that a complex network composed of 1,010 genes is organized to achieve few dominant modes demonstrates the utility of the systems approach for consolidating large amounts of genome-scale molecular information about a genome and its regulation to elucidate an organism's preferred environments and functional capabilities.

Bacterial Proteins↗