Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome scale models”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Inferring metabolic objectives and trade-offs in single cells during embryogenesis.

While proliferating cells optimize their metabolism to produce biomass, the metabolic objectives of cells that perform non-proliferative tasks are unclear. The opposing requirements for optimizing each objective result in a trade-off that forces single cells to prioritize their metabolic needs and optimally allocate limited resources. Here, we present single-cell optimization objective and trade-off inference (SCOOTI), which infers metabolic objectives and trade-offs in biological systems by integrating bulk and single-cell omics data, using metabolic modeling and machine learning. We validated SCOOTI by identifying essential genes from CRISPR-Cas9 screens in embryonic stem cells, and by inferring the metabolic objectives of quiescent cells, during different cell-cycle phases. Applying this to embryonic cell states, we observed a decrease in metabolic entropy upon development. We further uncovered a trade-off between glutathione and biosynthetic precursors in one-cell zygote, two-cell embryo, and blastocyst cells, potentially representing a trade-off between pluripotency and proliferation. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis↗

NAViFluX: a visualization‑centric platform for interactive analysis, refinement and design of genome‑scale metabolic networks.

MOTIVATION: Genome-scale metabolic network (GSMN) models enable flux-based metabolite fate discovery, metabolic engineering, drug target identification, and multi-omics integration. However, programming requirements, architectural complexity, and limited visualization support impede its adoption by the broader scientific community. Existing tools exclusively specialize in GSMN analyses or visualization while lacking important features such as pathway-specific views, database-integrated refinement, and comprehensive enrichment and perturbation analyses. RESULTS: Here, we present NAViFluX (metabolic Network Analysis and Visualization of Flux), a visualization-centric, web browser-based tool that unifies native pathway/subsystem map generation, interactive model refinement via KEGG/BiGG, pathway merging and modules for flux computations, topology, and functional enrichment all within network views. Using three independent case studies on Escherichia coli, the utility of NAViFluX for characterization of nutrient-specific metabolic adaptations, enhancing gene essentiality predictions and interpretability, and rational design of an optimized carbon-fixing metabolic state is demonstrated. AVAILABILITY AND IMPLEMENTATION: All source code and supplementary files associated with the case studies are publicly available via Zenodo at https://zenodo.org/records/19107831. NAViFluX can be easily installed as a standalone software through https://github.com/bnsb-lab-iith/NAViFluX.

Metabolic Networks and Pathways↗

Harnessing deep learning for proteome-scale detection of amyloid signaling motifs.

MOTIVATION: Amyloid signaling sequences adopt the cross-β fold that is capable of self-replication in the templating process. Propagation of the amyloid fold from the receptor to the effector protein is used for signal transduction in the immune response pathways in animals, fungi, and bacteria. So far, a dozen of families of amyloid signaling motifs (ASMs) have been classified. Unfortunately, due to the wide variety of ASMs it is difficult to identify them in large protein databases available, which limits the possibility of conducting experimental studies. To date, various deep learning (DL) models have been applied across a range of protein-related tasks, including domain family classification and the prediction of protein structure and protein-protein interactions. RESULTS: In this study, we develop tailor-made bidirectional LSTM and BERT-based architectures to model ASM, and compare their performance against a state-of-the-art machine learning grammatical model. Our research is focused on developing a discriminative model of generalized ASMs, capable of detecting ASMs in large datasets. The DL-based models are trained on a diverse set of motif families and a global negative set, and used to identify ASMs from remotely related families. We analyze how both models represent the data and demonstrate that the DL-based approaches effectively detect ASMs, including novel motifs, even at the genome scale. AVAILABILITY AND IMPLEMENTATION: The models are provided as a Python package, asmscan-bilstm, and a Docker image at https://github.com/chrispysz/asmscan-proteinbert-run. The source code can be accessed at https://github.com/jakub-galazka/asmscan-bilstm and https://github.com/chrispysz/asmscan-proteinbert. Data and results are at https://github.com/wdyrka-pwr/ASMscan.

Deep Learning↗

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines↗

Analysis of expressed sequence tags (ESTs) in the ciliated protozoan Tetrahymena thermophila.

To assess the utility of expressed sequence tag (EST) sequencing as a method of gene discovery in the ciliated protozoan Tetrahymena thermophila, we have sequenced either the 5' or 3' ends of 157 clones chosen at random from two cDNA libraries constructed from the mRNA of vegetatively growing cultures. Of 116 total non-redundant clones, 8.6% represented genes previously cloned in Tetrahymena. Fifty-two percent had significant identity to genes from other organisms represented in GenBank, of which 92% matched human proteins. Intriguing matches include an opioid-regulated protein, a glutamate-binding protein for an NMDA-receptor, and a stem-cell maintenance protein. Eleven-percent of the non-Tetrahymena specific matches were to genes present in humans and other mammals but not found in other model unicellular eukaryotes, including the completely sequenced Saccharomyces cerevisiae. Our data reinforce the fact that Tetrahymena is an excellent unicellular model system for studying many aspects of animal biology and is poised to become an important model system for genome-scale gene discovery and functional analysis.

Animals↗

Phylogenetic reconstruction of orthology, paralogy, and conserved synteny for dog and human.

Accurate predictions of orthology and paralogy relationships are necessary to infer human molecular function from experiments in model organisms. Previous genome-scale approaches to predicting these relationships have been limited by their use of protein similarity and their failure to take into account multiple splicing events and gene prediction errors. We have developed PhyOP, a new phylogenetic orthology prediction pipeline based on synonymous rate estimates, which accurately predicts orthology and paralogy relationships for transcripts, genes, exons, or genomic segments between closely related genomes. We were able to identify orthologue relationships to human genes for 93% of all dog genes from Ensembl. Among 1:1 orthologues, the alignments covered a median of 97.4% of protein sequences, and 92% of orthologues shared essentially identical gene structures. PhyOP accurately recapitulated genomic maps of conserved synteny. Benchmarking against predictions from Ensembl and Inparanoid showed that PhyOP is more accurate, especially in its predictions of paralogy. Nearly half (46%) of PhyOP paralogy predictions are unique. Using PhyOP to investigate orthologues and paralogues in the human and dog genomes, we found that the human assembly contains 3-fold more gene duplications than the dog. Species-specific duplicate genes, or "in-paralogues," are generally shorter and have fewer exons than 1:1 orthologues, which is consistent with selective constraints and mutation biases based on the sizes of duplicated genes. In-paralogues have experienced elevated amino acid and synonymous nucleotide substitution rates. Duplicates possess similar biological functions for either the dog or human lineages. Having accounted for 2,954 likely pseudogenes and gene fragments, and after separating 346 erroneously merged genes, we estimated that the human genome encodes a minimum of 19,700 protein-coding genes, similar to the gene count of nematode worms. PhyOP is a fast and robust approach to orthology prediction that will be applicable to whole genomes from multiple closely related species. PhyOP will be particularly useful in predicting orthology for mammalian genomes that have been incompletely sequenced, and for large families of rapidly duplicating genes.

Animals↗

Automated protein structure homology modeling: a progress report.

Understanding the molecular function of proteins is greatly enhanced by insights gained from their three-dimensional structures. Since experimental structures are only available for a small fraction of proteins, computational methods for protein structure modeling play an increasingly important role. Comparative protein structure modeling is currently the most accurate method, yielding models suitable for a wide spectrum of applications, such as structure-guided drug development or virtual screening. Stable and reliable automated prediction pipelines have been developed to apply large-scale comparative modeling to whole genomes or entire sequence databases. Model repositories give access to these annotated and evaluated models. In this review, we will discuss recent developments in automated comparative modeling and provide selected examples illustrating the use of homology models.

Animals↗

Evidence that rice and other cereals are ancient aneuploids.

Detailed analyses of the genomes of several model organisms revealed that large-scale gene or even entire-genome duplications have played prominent roles in the evolutionary history of many eukaryotes. Recently, strong evidence has been presented that the genomic structure of the dicotyledonous model plant species Arabidopsis is the result of multiple rounds of entire-genome duplications. Here, we analyze the genome of the monocotyledonous model plant species rice, for which a draft of the genomic sequence was published recently. We show that a substantial fraction of all rice genes ( approximately 15%) are found in duplicated segments. Dating of these block duplications, their nonuniform distribution over the different rice chromosomes, and comparison with the duplication history of Arabidopsis suggest that rice is not an ancient polyploid, as suggested previously, but an ancient aneuploid that has experienced the duplication of one-or a large part of one-chromosome in its evolutionary past, approximately 70 million years ago. This date predates the divergence of most of the cereals, and relative dating by phylogenetic analysis shows that this duplication event is shared by most if not all of them.

Aneuploidy↗

Systems biology as a foundation for genome-scale synthetic biology.

As the ambitions of synthetic biology approach genome-scale engineering, comprehensive characterization of cellular systems is required, as well as a means to accurately model cell-scale molecular interactions. These requirements are coincident with the goals of systems biology and, thus, systems biology will become the foundation for genome-scale synthetic biology. Systems biology will form this foundation through its efforts to reconstruct and integrate cellular systems, develop the mathematics, theory and software tools for the accurate modeling of these integrated systems, and through evolutionary mechanisms. As genome-scale synthetic biology is so enabled, it will prove to be a positive feedback driver of systems biology by exposing and forcing researchers to confront those aspects of systems biology which are inadequately understood.

Biology↗

Scaling linear-model breeding values to the liability scale: an application to pig binary traits.

In commercial pig production, many important traits are recorded as binary phenotypes. For such traits, threshold models offer an appropriate framework but are computationally intensive. Thus, linear models are widely used to obtain genomic estimated breeding values (GEBV); however, these are on the observed scale (phenotypic). This creates the need for a robust method to approximate GEBV from linear models to the liability scale. A recently proposed approximation showed good concordance for low-prevalence traits (<5%) but has not yet been tested for a wider range of prevalence values and for models with more than one random effect. We aimed to evaluate the performance of this approximation for pig binary traits with prevalences ranging from <5% to >86%, in both animal and maternal animal models. Data were available for five fitness traits (FT1-FT5), with up to 233k animals with phenotypes, of which 204k animals were genotyped with a 25k SNP array. Variance component estimates were obtained using threshold models. Classical animal models were used for FT1-FT3, and maternal animal models for FT4 and FT5. Variance components on the observed scale were then obtained by multiplying estimates from a threshold model by the square of the height of the standard normal density evaluated at the threshold. GEBV were predicted using single-step genomic best linear unbiased prediction under both linear and threshold models. The approximation tested involved scaling the GEBV using the height of the ordinate of the standard normal distribution evaluated at the threshold as a scaling factor. The agreement between GEBV from the scaled linear model and the threshold model on the probability scale was evaluated using Pearson and Spearman correlations, mean squared error (MSE), regression parameters, overlapping coefficient (OVL), distribution overlap, and classification accuracy (CACC). Correlations between linear and threshold GEBV ranged from 0.94 (low-prevalence traits) to 0.99 (high-prevalence traits) for the direct GEBV and were 0.99 for the maternal GEBV. MSE were close to zero. The OVL exceeded 0.83 for all traits. CACC ranged from 95.10% to 98.33% for the direct GEBV and from 92.54% to 97.42% for the maternal GEBV. Regardless of model and trait prevalence, this approximation yielded GEBV that are highly consistent with threshold model GEBV, providing a reliable, practical approach for large-scale pig genetic evaluations for binary traits using linear models.

Animals↗

Gompertz mortality law and scaling behavior of the Penna model.

The Penna model is a model of evolutionary ageing through mutation accumulation where traditionally time and the age of an organism are treated as discrete variables and an organism's genome is represented by a binary bit string. We reformulate the asexual Penna model and show that a universal scale invariance emerges as we increase the number of discrete genome bits to the limit of a continuum. The continuum model, introduced by Almeida and Thomas [Int. J. Mod. Phys. C 11, 1209 (2000)] can be recovered from the discrete model in the limit of infinite bits coupled with a vanishing mutation rate per bit. Finally, we show that scale invariant properties may lead to the ubiquitous Gompertz law for mortality rates for early ages, which is generally regarded as being empirical.

Aging↗

Modeling gene and genome duplications in eukaryotes.

Recent analysis of complete eukaryotic genome sequences has revealed that gene duplication has been rampant. Moreover, next to a continuous mode of gene duplication, in many eukaryotic organisms the complete genome has been duplicated in their evolutionary past. Such large-scale gene duplication events have been associated with important evolutionary transitions or major leaps in development and adaptive radiations of species. Here, we present an evolutionary model that simulates the duplication dynamics of genes, considering genome-wide duplication events and a continuous mode of gene duplication. Modeling the evolution of the different functional categories of genes assesses the importance of different duplication events for gene families involved in specific functions or processes. By applying our model to the Arabidopsis genome, for which there is compelling evidence for three whole-genome duplications, we show that gene loss is strikingly different for large-scale and small-scale duplication events and highly biased toward certain functional classes. We provide evidence that some categories of genes were almost exclusively expanded through large-scale gene duplication events. In particular, we show that the three whole-genome duplications in Arabidopsis have been directly responsible for >90% of the increase in transcription factors, signal transducers, and developmental genes in the last 350 million years. Our evolutionary model is widely applicable and can be used to evaluate different assumptions regarding small- or large-scale gene duplication events in eukaryotic genomes.

Arabidopsis↗

Toward large-scale modeling of the microbial cell for computer simulation.

In the post-genomic era, the large-scale, systematic, and functional analysis of all cellular components using transcriptomics, proteomics, and metabolomics, together with bioinformatics for the analysis of the massive amount of data generated by these "omics" methods are the focus of intensive research activities. As a consequence of these developments, systems biology, whose goal is to comprehend the organism as a complex system arising from interactions between its multiple elements, becomes a more tangible objective. Mathematical modeling of microorganisms and subsequent computer simulations are effective tools for systems biology, which will lead to a better understanding of the microbial cell and will have immense ramifications for biological, medical, environmental sciences, and the pharmaceutical industry. In this review, we describe various types of mathematical models (structured, unstructured, static, dynamic, etc.), of microorganisms that have been in use for a while, and others that are emerging. Several biochemical/cellular simulation platforms to manipulate such models are summarized and the E-Cell system developed in our laboratory is introduced. Finally, our strategy for building a "whole cell metabolism model", including the experimental approach, is presented.

Biotechnology↗

Genotype by Environment Interactions in Gene Regulation Underlie the Response to Soil Drying in the Model Grass Brachypodium distachyon.

Gene expression is a quantitative trait under the control of genetic and environmental factors and their interaction, so-called genotype and environment (G &#xd7; E). Understanding the mechanisms driving G &#xd7; E is fundamental for ensuring stable crop performance across environments and for predicting the response of natural populations to climate change. Gene expression is regulated through complex molecular networks, yet the interactions between genotype and environment in gene regulation are rarely considered, particularly at the genome scale. Current frameworks and experimental designs often lack power to explicitly test network rewiring or to systematically compare regulatory networks. Here, we leverage a highly replicated RNA-sequencing dataset to model genome-scale gene expression variation between two natural accessions of the model grass Brachypodium distachyon and their response to soil drying. We first identified genotypic, environmental, and G &#xd7; E effects on physiological, metabolic, and gene expression traits. We identify patterns of conservation-or variation-in gene coexpression networks and link these coexpression features to physiological traits. We further develop predictions of gene-gene interactions using causal inference and screen for interactions specific to-or with higher affinity in-a single genotype, treatment, or their interaction, G &#xd7; E. Our analyses identify variation in candidate gene regulatory networks that may shape the evolution of environmental response in B. distachyon. We highlight the environmentally dependent regulatory control of several metabolic traits shown previously to play a role in drought acclimation. The framework presented here provides a scalable approach for more complex comparisons, particularly with the growing availability of large datasets from technologies such as single-cell transcriptomics.

Brachypodium↗

Dynamic simulation of protein complex formation on a genomic scale.

MOTIVATION: One of the central questions in the post-genomic era is the understanding of protein-protein interactions and of protein complex formation. It has been observed that protein complex size distributions of the yeast Saccharomyces cerevisiae decay exponentially. The shape of these size distributions reflects mechanisms of protein complex association and dissociation. RESULTS: We present the most simple dynamic model that is able to reproduce the observed protein complex size distribution for yeast. This protein association-dissociation model (PAD-model) simulates the dynamics of protein complex formation on a genomic scale for about 50 million protein molecules. By ruling out different model variants it is possible to elucidate fundamental features of the protein complex dynamics, e.g. complex association is independent of complex size. In addition, the PAD-model provides information about the complexity of the yeast proteome and it gives an idea of how many complexes could not be identified during the measurements. AVAILABILITY: All programs used for this publication are available on request from the authors. CONTACT: beyer@imb-jena.de SUPPLEMENTARY INFORMATION: Supplementary information about the model and its interpretation can be downloaded from http://www.imb-jena.de/tsb/pad.

Computer Simulation↗

An object model and database for functional genomics.

MOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site.

Abstracting and Indexing↗

Using metabolic flux data to further constrain the metabolic solution space and predict internal flux patterns: the Escherichia coli spectrum.

Constraint-based metabolic modeling has been used to capture the genome-scale, systems properties of an organism's metabolism. The first generation of these models has been built on annotated gene sequence. To further this field, we now need to develop methods to incorporate additional "omic" data types including transcriptomics, metabolomics, and fluxomics to further facilitate the construction, validation, and predictive capabilities of these models. The work herein combines metabolic flux data with an in silico model of central metabolism of Escherichia coli for model centric integration of the flux data. The extreme pathways for this network, which define the allowable solution space for all possible flux distributions, are analyzed using the alpha-spectrum. The alpha-spectrum determines which extreme pathways can and cannot contribute to the metabolic flux distribution for a given condition and gives the allowable range of weightings on each extreme pathway that can contribute. Since many extreme pathways cannot be used under certain conditions, the result is a "condition-specific" solution space that is a subset of the original solution space. The alpha-spectrum results are used to create a "condition-specific" extreme pathway matrix that can be analyzed using singular value decomposition (SVD). The first mode of the SVD analysis characterizes the solution space for a given condition. We show that SVD analysis of the alpha-spectrum extreme pathway matrix that incorporates measured uptake and byproduct secretion rates, can predict internal flux trends for different experimental conditions. These predicted internal flux trends are, in general, consistent with the flux trends measured using experimental metabolic flux analysis techniques.

Ammonia↗

From DNA sequence analysis to modeling replication in the human genome.

We explore the large-scale behavior of nucleotide compositional strand asymmetries along human chromosomes. As we observe for 7 of 9 origins of replication experimentally identified so far, the (TA+GC) skew displays rather sharp upward jumps, with a linear decreasing profile in between two successive jumps. We present a model of replication with well positioned replication origins and random terminations that accounts for the observed characteristic serrated skew profiles. We succeed in identifying 287 pairs of putative adjacent replication origins with an origin spacing approximately 1-2 Mbp that are likely to correspond to replication foci observed in interphase nuclei and recognized as stable structures that persist throughout subsequent cell generations.

DNA Replication↗