Search PubMed⌕ Search

Biomedical subjects

Eberhard O Voit

Publications and source records attributed to Eberhard O Voit.

At least 19 recordsLinked to original sources

Parameter estimation in biochemical systems models with alternating regression.

BACKGROUND: The estimation of parameter values continues to be the bottleneck of the computational analysis of biological systems. It is therefore necessary to develop improved methods that are effective, fast, and scalable. RESULTS: We show here that alternating regression (AR), applied to S-system models and combined with methods for decoupling systems of differential equations, provides a fast new tool for identifying parameter values from time series data. The key feature of AR is that it dissects the nonlinear inverse problem of estimating parameter values into iterative steps of linear regression. We show with several artificial examples that the method works well in many cases. In cases of no convergence, it is feasible to dedicate some computational effort to identifying suitable start values and search settings, because the method is fast in comparison to conventional methods that the search for suitable initial values is easily recouped. Because parameter estimation and the identification of system structure are closely related in S-system modeling, the AR method is beneficial for the latter as well. Specifically, we show with an example from the literature that AR is three to five orders of magnitudes faster than direct structure identifications in systems of nonlinear differential equations. CONCLUSION: Alternating regression provides a strategy for the estimation of parameter values and the identification of structure and regulation in S-systems that is genuinely different from all existing methods. Alternating regression is usually very fast, but its convergence patterns are complex and will require further investigation. In cases where convergence is an issue, the enormous speed of the method renders it feasible to select several initial guesses and search settings as an effective countermeasure.

Computational Biology↗

A multivariate prediction model for microarray cross-hybridization.

BACKGROUND: Expression microarray analysis is one of the most popular molecular diagnostic techniques in the post-genomic era. However, this technique faces the fundamental problem of potential cross-hybridization. This is a pervasive problem for both oligonucleotide and cDNA microarrays; it is considered particularly problematic for the latter. No comprehensive multivariate predictive modeling has been performed to understand how multiple variables contribute to (cross-) hybridization. RESULTS: We propose a systematic search strategy using multiple multivariate models [multiple linear regressions, regression trees, and artificial neural network analyses (ANNs)] to select an effective set of predictors for hybridization. We validate this approach on a set of DNA microarrays with cytochrome p450 family genes. The performance of our multiple multivariate models is compared with that of a recently proposed third-order polynomial regression method that uses percent identity as the sole predictor. All multivariate models agree that the 'most contiguous base pairs between probe and target sequences,' rather than percent identity, is the best univariate predictor. The predictive power is improved by inclusion of additional nonlinear effects, in particular target GC content, when regression trees or ANNs are used. CONCLUSION: A systematic multivariate approach is provided to assess the importance of multiple sequence features for hybridization and of relationships among these features. This approach can easily be applied to larger datasets. This will allow future developments of generalized hybridization models that will be able to correct for false-positive cross-hybridization signals in expression experiments.

Algorithms↗

Identification of metabolic system parameters using global optimization methods.

BACKGROUND: The problem of estimating the parameters of dynamic models of complex biological systems from time series data is becoming increasingly important. METHODS AND RESULTS: Particular consideration is given to metabolic systems that are formulated as Generalized Mass Action (GMA) models. The estimation problem is posed as a global optimization task, for which novel techniques can be applied to determine the best set of parameter values given the measured responses of the biological system. The challenge is that this task is nonconvex. Nonetheless, deterministic optimization techniques can be used to find a global solution that best reconciles the model parameters and measurements. Specifically, the paper employs branch-and-bound principles to identify the best set of model parameters from observed time course data and illustrates this method with an existing model of the fermentation pathway in Saccharomyces cerevisiae. This is a relatively simple yet representative system with five dependent states and a total of 19 unknown parameters of which the values are to be determined. CONCLUSION: The efficacy of the branch-and-reduce algorithm is illustrated by the S. cerevisiae example. The method described in this paper is likely to be widely applicable in the dynamic modeling of metabolic networks.

Algorithms↗

An automated procedure for the extraction of metabolic network information from time series data.

Novel high-throughput measurement techniques in vivo are beginning to produce dense high-quality time series which can be used to investigate the structure and regulation of biochemical networks. We propose an automated information extraction procedure which takes advantage of the unique S-system structure and supports model building from time traces, curve fitting, model selection, and structure identification based on parameter estimation. The procedure comprises of three modules: model Generation, parameter estimation or model Fitting, and model Selection (GFS algorithm). The GFS algorithm has been implemented in MATLAB and returns a list of candidate S-systems which adequately explain the data and guides the search to the most plausible model for the time series under study. By combining two strategies (namely decoupling and limiting connectivity) with methods of data smoothing, the proposed algorithm is scalable up to realistic situations of moderate size. We illustrate the proposed methodology with a didactic example.

Algorithms↗

Parameter estimation in modulated, unbranched reaction chains within biochemical systems.

Modern biology is increasingly developing techniques for measuring time series of global gene expression and of many simultaneous proteins or metabolites. These data contain valuable information on the dynamics of cells, which has to be extracted with computational means. Given a suitable mathematical model, this extraction is in principle a straightforward regression task, but the complexity and nonlinearity of the differential equations that describe biological systems cause severe difficulties when the systems are of realistic size. We propose a method of stepwise regression that can be applied effectively to linear portions of pathways. The method may be combined with other estimation methods and either directly yields reasonable parameter estimates or at least provides appropriate start values for subsequent nonlinear search algorithms. We illustrate the method with the analysis of in vivo NMR data describing the dynamics of glycolytic metabolites in Lactococcus lactis.

Algorithms↗

Computation and analysis of time-dependent sensitivities in Generalized Mass Action systems.

Understanding biochemical system dynamics is becoming increasingly important for insights into the functioning of organisms and for biotechnological manipulations, and additional techniques and methods are needed to facilitate investigations of dynamical properties of systems. Extensions to the method of Ingalls and Sauro, addressing time-dependent sensitivity analysis, provide a new tool for executing such investigations. We present here the results of sample analyses using time-dependent sensitivities for three model systems taken from the literature, namely an anaerobic fermentation pathway in yeast, a negative feedback oscillator modeling cell-cycle phenomena, and the Mitogen Activated Protein (MAP) kinase cascade. The power of time-dependent sensitivities is particularly evident in the case of the MAPK cascade. In this example it is possible to identify the emergence of a concentration of MAPKK that provides the best response with respect to rapid and efficient activation of the cascade, while over- and under-expression of MAPKK relative to this concentration have qualitatively different effects on the transient response of the cascade. Also of interest is the quite general observation that phase-plane representations of sensitivities in oscillating systems provide insights into the manner with which perturbations in the envelope of the oscillation result from small changes in initial concentrations of components of the oscillator. In addition to these applied analyses, we present an algorithm for the efficient computation of time-dependent sensitivities for Generalized Mass Action (GMA) systems, the most general of the canonical system representations of Biochemical Systems Theory (BST). The algorithm is shown to be comparable to, or better than, other methods of solution, as exemplified with three biochemical systems taken from the literature.

Animals↗

Simulation and validation of modelled sphingolipid metabolism in Saccharomyces cerevisiae.

Mathematical models have become a necessary tool for organizing the rapidly increasing amounts of large-scale data on biochemical pathways and for advanced evaluation of their structure and regulation. Most of these models have addressed specific pathways using either stoichiometric or flux-balance analysis, or fully kinetic Michaelis-Menten representations, metabolic control analysis, or biochemical systems theory. So far, the predictions of kinetic models have rarely been tested using direct experimentation. Here, we validate experimentally a biochemical systems theoretical model of sphingolipid metabolism in yeast. Simulations of metabolic fluxes, enzyme deletion and the effects of inositol (a key regulator of phospholipid metabolism) led to predictions that show significant concordance with experimental results generated post hoc. The model also allowed the simulation of the effects of acute perturbations in fatty-acid precursors of sphingolipids, a situation that is not amenable to direct experimentation. The results demonstrate that modelling now allows testable predictions as well as the design and evaluation of hypothetical 'thought experiments' that may generate new metabolomic approaches.

Computer Simulation↗

Controllability of non-linear biochemical systems.

Mathematical methods of biochemical pathway analysis are rapidly maturing to a point where it is possible to provide objective rationale for the natural design of metabolic systems and where it is becoming feasible to manipulate these systems based on model predictions, for instance, with the goal of optimizing the yield of a desired microbial product. So far, theory-based metabolic optimization techniques have mostly been applied to steady-state conditions or the minimization of transition time, using either linear stoichiometric models or fully kinetic models within biochemical systems theory (BST). This article addresses the related problem of controllability, where the task is to steer a non-linear biochemical system, within a given time period, from an initial state to some target state, which may or may not be a steady state. For this purpose, BST models in S-system form are transformed into affine non-linear control systems, which are subjected to an exact feedback linearization that permits controllability through independent variables. The method is exemplified with a small glycolytic-glycogenolytic pathway that had been analyzed previously by several other authors in different contexts.

Biochemical Phenomena↗

Challenges for the identification of biological systems from in vivo time series data.

Modern methods of high-throughput molecular biology render it possible to generate time series of metabolite concentrations and the expression of genes and proteins in vivo. These time profiles contain valuable information about the structure and dynamics of the underlying biological system. This information is implicit and its extraction is a challenging but ultimately very rewarding task for the mathematical modeler. Using a well-suited modeling framework, such as Biochemical Systems Theory (BST), it is possible to formulate the extraction of information as an inverse problem that in principle may be solved with a genetic algorithm or nonlinear regression. However, two types of issues associated with this inverse problem make the extraction task difficult. One type pertains to the algorithmic difficulties encountered in nonlinear regressions with moderate and large systems. The other type is of an entirely different nature. It is a consequence of assumptions that are often taken for granted in the design and analysis of mathematical models of biological systems and that need to be revisited in the context of inverse analyses. The article describes the extraction process and some of its challenges and proposes partial solutions.

Linear Models↗

Priming nonlinear searches for pathway identification.

BACKGROUND: Dense time series of metabolite concentrations or of the expression patterns of proteins may be available in the near future as a result of the rapid development of novel, high-throughput experimental techniques. Such time series implicitly contain valuable information about the connectivity and regulatory structure of the underlying metabolic or proteomic networks. The extraction of this information is a challenging task because it usually requires nonlinear estimation methods that involve iterative search algorithms. Priming these algorithms with high-quality initial guesses can greatly accelerate the search process. In this article, we propose to obtain such guesses by preprocessing the temporal profile data and fitting them preliminarily by multivariate linear regression. RESULTS: The results of a small-scale analysis indicate that the regression coefficients reflect the connectivity of the network quite well. Using the mathematical modeling framework of Biochemical Systems Theory (BST), we also show that the regression coefficients may be translated into constraints on the parameter values of the nonlinear BST model, thereby reducing the parameter search space considerably. CONCLUSION: The proposed method provides a good approach for obtaining a preliminary network structure from dense time series. This will be more valuable as the systems become larger, because preprocessing and effective priming can significantly limit the search space of parameters defining the network connectivity, thereby facilitating the nonlinear estimation task.

Models, Biological↗

Improved methods for the mathematically controlled comparison of biochemical systems.

The method of mathematically controlled comparison provides a structured approach for the comparison of alternative biochemical pathways with respect to selected functional effectiveness measures. Under this approach, alternative implementations of a biochemical pathway are modeled mathematically, forced to be equivalent through the application of selected constraints, and compared with respect to selected functional effectiveness measures. While the method has been applied successfully in a variety of studies, we offer recommendations for improvements to the method that (1) relax requirements for definition of constraints sufficient to remove all degrees of freedom in forming the equivalent alternative, (2) facilitate generalization of the results thus avoiding the need to condition those findings on the selected constraints, and (3) provide additional insights into the effect of selected constraints on the functional effectiveness measures. We present improvements to the method and related statistical models, apply the method to a previously conducted comparison of network regulation in the immune system, and compare our results to those previously reported.

Animals↗

Decoupling dynamical systems for pathway identification from metabolic profiles.

RATIONALE: Modern molecular biology is generating data of unprecedented quantity and quality. Particularly exciting for biochemical pathway modeling and proteomics are comprehensive, time-dense profiles of metabolites or proteins that are measurable, for instance, with mass spectrometry, nuclear magnetic resonance or protein kinase phosphorylation. These profiles contain a wealth of information about the structure and dynamics of the pathway or network from which the data were obtained. The retrieval of this information requires a combination of computational methods and mathematical models, which are typically represented as systems of ordinary differential equations. RESULTS: We show that, for the purpose of structure identification, the substitution of differentials with estimated slopes in non-linear network models reduces the coupled system of differential equations to several sets of decoupled algebraic equations, which can be processed efficiently in parallel or sequentially. The estimation of slopes for each time series of the metabolic or proteomic profile is accomplished with a 'universal function' that is computed directly from the data by cross-validated training of an artificial neural network (ANN). CONCLUSIONS: Without preprocessing, the inverse problem of determining structure from metabolic or proteomic profile data is challenging and computationally expensive. The combination of system decoupling and data fitting with universal functions simplifies this inverse problem very significantly. Examples show successful estimations and current limitations of the method. AVAILABILITY: A preliminary Web-based application for ANN smoothing is accessible at http://bioinformatics.musc.edu/webmetabol/. S-systems can be interactively analyzed with the user-friendly freeware PLAS (http://correio.cc.fc.ul.pt/~aenf/plas.html) or with the MATLAB module BSTLab (http://bioinformatics.musc.edu/bstlab/), which is currently being beta-tested.

Algorithms↗

Integration of kinetic information on yeast sphingolipid metabolism in dynamical pathway models.

For the first time, kinetic information from the literature was collected and used to construct integrative dynamical mathematical models of sphingolipid metabolism. One model was designed primarily with kinetic equations in the tradition of Michaelis and Menten whereas the other two models were designed as alternative power-law models within the framework of Biochemical Systems Theory. Each model contains about 50 variables, about a quarter of which are dependent (state) variables, while the others are independent inputs and enzyme activities that are considered constant. The models account for known regulatory signals that exert control over the pathway. Standard mathematical testing, repeated revisiting of the literature, and numerous rounds of amendments and refinements resulted in models that are stable and rather insensitive to perturbations in inputs or parameter values. The models also appear to be compatible with the modest amount of experimental experience that lends itself to direct comparisons. Even though the three models are based on different mathematical representations, they show dynamic responses to a variety of perturbations and changes in conditions that are essentially equivalent for small perturbations and similar for large perturbations. The kinetic information used for model construction and the models themselves can serve as a starting point for future analyses and refinements.

Models, Biological↗

Analysis of dynamic labeling data.

Comprehensive assessments of the organization and regulation of metabolic pathways cannot be limited to steady-state measurements alone but require dynamic time series data. One experimental means of generating such data consists of radioactively labeling precursors and measuring their fate over time. While labeling experiments belong to the standard repertoire of biological laboratory techniques, corresponding mathematical tools for analyzing the non-linear dynamics of tracers are scarce. The article addresses this issue, using Biochemical Systems Theory as the modeling framework. The description of the dynamics of labeled metabolites alone is difficult, but it is demonstrated that these difficulties are easily overcome by setting up dynamic models in two or three blocks, one for the kinetics of the total pools, the second just for the labeled portions, and the third, optional, block for the remaining unlabeled components. Since the dynamic model is not limited in complexity and can account for linear pathways, converging and diverging branches, cycles, and the various observed modes of regulation, the proposed method of non-linear tracer analysis is rather general and permits simulations of most standard labeling experiments, both at steady state and during transients.

Computer Simulation↗

Yeast sphingolipid metabolism: clues and connections.

This review of sphingolipid metabolism in the budding yeast Saccharomyces cerevisiae contains information on the enzymes and the genes that encode them, as well as connections to other metabolic pathways. Particular attention is given to yeast homologs, domains, and motifs in the sequence, cellular localization of enzymes, and possible protein-protein interactions. Also included are genetic interactions of special interest that provide clues to the cellular biological roles of particular sphingolipid metabolic pathways and specific sphingolipids.

Models, Biological↗

A quantitative model of the generation of N(epsilon)-(carboxymethyl)lysine in the Maillard reaction between collagen and glucose.

The Maillard reaction between reducing sugars and amino groups of biomolecules generates complex structures known as AGEs (advanced glycation endproducts). These have been linked to protein modifications found during aging, diabetes and various amyloidoses. To investigate the contribution of alternative routes to the formation of AGEs, we developed a mathematical model that describes the generation of CML [ N(epsilon)-(carboxymethyl)lysine] in the Maillard reaction between glucose and collagen. Parameter values were obtained by fitting published data from kinetic experiments of Amadori compound decomposition and glycoxidation of collagen by glucose. These raw parameter values were subsequently fine-tuned with adjustment factors that were deduced from dynamic experiments taking into account the glucose and phosphate buffer concentrations. The fine-tuned model was used to assess the relative contributions of the reaction between glyoxal and lysine, the Namiki pathway, and Amadori compound degradation to the generation of CML. The model suggests that the glyoxal route dominates, except at low phosphate and high glucose concentrations. The contribution of Amadori oxidation is generally the least significant at low glucose concentrations. Simulations of the inhibition of CML generation by aminoguanidine show that this compound effectively blocks the glyoxal route at low glucose concentrations (5 mM). Model results are compared with literature estimates of the contributions to CML generation by the three pathways. The significance of the dominance of the glyoxal route is discussed in the context of possible natural defensive mechanisms and pharmacological interventions with the goal of inhibiting the Maillard reaction in vivo.

Collagen↗

Extending knowledge of Escherichia coli metabolism by modeling and experiment.

One of the challenges for 'post-genomic' biology is the integration of data from many different sources. Two recent studies independently take steps towards this goal for Escherichia coli, using mathematical modeling and a combination of gene expression and protein levels to predict new gene functions and metabolic behaviors.

Escherichia coli↗

Biochemical and genomic regulation of the trehalose cycle in yeast: review of observations and canonical model analysis.

The physiological hallmark of heat-shock response in yeast is a rapid, enormous increase in the concentration of trehalose. Normally found in growing yeast cells and other organisms only as traces, trehalose becomes a crucial protector of proteins and membranes against a variety of stresses, including heat, cold, starvation, desiccation, osmotic or oxidative stress, and exposure to toxicants. Trehalose is produced from glucose 6-phosphate and uridine diphosphate glucose in a two-step process, and recycled to glucose by trehalases. Even though the trehalose cycle consists of only a few metabolites and enzymatic steps, its regulatory structure and operation are surprisingly complex. The article begins with a review of experimental observations on the regulation of the trehalose cycle in yeast and proposes a canonical model for its analysis. The first part of this analysis demonstrates the benefits of the various regulatory features by means of controlled comparisons with models of otherwise equivalent pathways lacking these features. The second part elucidates the significance of the expression pattern of the trehalose cycle genes in response to heat shock. Interestingly, the genes contributing to trehalose formation are up-regulated to very different degrees, and even the trehalose degrading trehalases show drastically increased activity during heat-shock response. Again using the method of controlled comparisons, the model provides rationale for the observed pattern of gene expression and reveals benefits of the counterintuitive trehalase up-regulation.

Feedback, Physiological↗