Statistical methods and microarray data.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Andrei Yakovlev.
Explore the source record for details and available documents.
This paper considers the utility of a new class of stochastic branching processes with non-homogeneous immigration in modeling complex renewing cell systems. Such systems typically include the population of stem cells that provides an inexhaustible supply of cells necessary for maintaining the cellular composition of a tissue. A stem cell may be induced to transform (differentiate) into a progenitor cell. Progenitor cells retain the ability to proliferate and their function is believed to provide a quick proliferative response to an increased demand for cells in the population. There may be several sub-types of progenitor cells. Terminally differentiated cells do not divide under normal conditions; they are responsible for maintaining tissue-specific functions. Recent advancements in experimental techniques offer considerable scope for quantitative studies of in vivo cell kinetics based on stochastic modeling of renewing cell populations. However, no ready-made theory is currently available to take full advantage of these advancements. This paper introduces such a theory with a special focus on its feasibility in biological applications.
BACKGROUND: The purpose of this paper is two-fold. The first objective is to validate the assumptions behind a stochastic model developed earlier by these authors to describe oligodendrocyte generation in cell culture. The second is to generate time-lapse data that may help biomathematicians to build stochastic models of cell proliferation and differentiation under other experimental scenarios. RESULTS: Using time-lapse video recording it is possible to follow the individual evolutions of different cells within each clone. This experimental technique is very laborious and cannot replace model-based quantitative inference from clonal data. However, it is unrivalled in validating the structure of a stochastic model intended to describe cell proliferation and differentiation at the clonal level. In this paper, such data are reported and analyzed for oligodendrocyte precursor cells cultured in vitro. CONCLUSION: The results strongly support the validity of the most basic assumptions underpinning the previously proposed model of oligodendrocyte development in cell culture. However, there are some discrepancies; the most important is that the contribution of progenitor cell death to cell kinetics in this experimental system has been underestimated.
One of the prevailing ideas in the literature on microarray data analysis is to pool the expression measures across genes and treat them as a sample drawn from some distribution. Several universal laws were proposed to analytically describe this distribution. This idea raises a number of concerns. The expression levels of genes are not identically distributed random variables so that treating them as a sample amounts to sampling from a mixture of equally weighted distributions, each being associated with a different gene. The expression levels of different genes are heavily dependent random variables so that the law of large numbers and statistical goodness-of-fit tests are normally inapplicable to this kind of data. This dependence represents a very serious pitfall in microarray data analysis.
Modern methods of microarray data analysis are biased towards selecting those genes that display the most pronounced differential expression. The magnitude of differential expression does not necessarily indicate biological significance and other criteria are needed to supplement the information on differential expression. Three large sets of microarray data on childhood leukemia were analyzed by an original method introduced in this paper. A new type of stochastic dependence between expression levels in gene pairs was deciphered by our analysis. This modulation-like unidirectional dependence between expression signals arises when the expression of a "gene-modulator'' is stochastically proportional to that of a "gene-driver''. A total of more than 35% of all pairs formed from 12550 genes were conservatively estimated to belong to this type. There are genes that tend to form Type A relationships with the overwhelming majority of genes. However, this picture is not static: the composition of Type A gene pairs may undergo dramatic changes when comparing two phenotypes. The ability to identify genes that act as ;;modulators'' provides a potential strategy of prioritizing candidate genes.
BACKGROUND: The number of genes declared differentially expressed is a random variable and its variability can be assessed by resampling techniques. Another important stability indicator is the frequency with which a given gene is selected across subsamples. We have conducted studies to assess stability and some other properties of several gene selection procedures with biological and simulated data. RESULTS: Using resampling techniques we have found that some genes are selected much less frequently (across sub-samples) than other genes with the same adjusted p-values. The extent to which this type of instability manifests itself can be assessed by a method introduced in this paper. The effect of correlation between gene expression levels on the performance of multiple testing procedures is studied by computer simulations. CONCLUSION: Resampling represents a tool for reducing the set of initially selected genes to those with a sufficiently high selection frequency. Using resampling techniques it is also possible to assess variability of different performance indicators. Stability properties of several multiple testing procedures are described at length in the present paper.
Some extended false discovery rate (FDR) controlling multiple testing procedures rely heavily on empirical estimates of the FDR constructed from gene expression data. Such estimates are also used as performance indicators when comparing different methods for microarray data analysis. The present communication shows that the variance of the proposed estimators may be intolerably high, the correlation structure of microarray data being the main cause of their instability.
Stochastic dependence between gene expression levels in microarray data is of critical importance for the methods of statistical inference that resort to pooling test statistics across genes. The empirical Bayes methodology in the nonparametric and parametric formulations, as well as closely related methods employing a two-component mixture model, represent typical examples. It is frequently assumed that dependence between gene expressions (or associated test statistics) is sufficiently weak to justify the application of such methods for selecting differentially expressed genes. By applying resampling techniques to simulated and real biological data sets, we have studied a potential impact of the correlation between gene expression levels on the statistical inference based on the empirical Bayes methodology. We report evidence from these analyses that this impact may be quite strong, leading to a high variance of the number of differentially expressed genes. This study also pinpoints specific components of the empirical Bayes method where the reported effect manifests itself.
A comprehensive entirely mechanistic model of the kinetics of cell population in vitro exposed to continuous irradiation is formulated and analysed. The model provides a stochastic description of the processes of formation and repair of radiation-induced lesions, as well as of cell cycling, cell proliferation, and cell death. Unobservable kinetic parameters of the model are estimated from experimental data on the sizes of clones formed in synchronized cultures of S3 HeLa cells exposed to continuous irradiation by gamma-rays at various dose rates. The effects of dose rate on cell cycle duration and damage repair kinetics are studied. Continuous and acute exposures with the same total dose are compared in terms of their impact on cell survival.
This paper presents a new method to analyze clonal data on oligodendrocyte development in cell culture. The process of oligodendrocyte generation from precursor cells is modelled as a multi-type Bellman-Harris branching process as suggested in an earlier paper [K. Boucher, A. Zorin, A.Y. Yakovlev, M. Mayer-Proschel, M. Noble, An alternative stochastic model of generation of oligodendrocytes in cell culture, J. Math. Biol. 43 (2001) 22]. This model has been extended to allow for death of oligodendrocytes as well as a dissimilar distribution of the first mitotic cycle duration as compared to the subsequent cycles of precursor cells, which lengths are assumed to be independent and identically distributed random variables. Since the time-span of oligodendrocytes is not directly observable in clonal data, plausible parametric assumptions are invoked to make estimation problems tractable. In particular, the time to cell death follows a two-parameter gamma distribution, while the lapse of time between the event of cell death and the event of cell disintegration is assumed to be exponentially distributed. A simulated pseudo maximum likelihood method for estimation of model parameters has been developed using simulation-based approximations of the expected numbers and variance-covariance matrices for different types of cells. Finite sample properties of the estimation procedure are studied by computer simulations. The proposed method is illustrated with an analysis of the clonal development of O-2A progenitor cells isolated from the rat optic nerve and the corpus callosum.
This article presents a stochastic model designed to analyze experimental data on the development of cell clones composed of two (or more) distinct types of cells. The proposed model is an extension of the traditional multi-type Bellman-Harris branching stochastic process allowing for nonidentical time-to-transformation distributions defined for different cell types. A simulated pseudo likelihood method has been developed for the parametric statistical inference from experimental data on cell clones under the proposed model. The method uses simulation-based approximations of the means and the variance-covariance matrices of cell counts. The proposed estimator for the vector of unknown parameters is strongly consistent and asymptotically normal under mild regularity conditions, while its variance-covariance matrix is estimated by the parametric bootstrap. A Monte Carlo Wald test is proposed for the test of hypotheses. Finite sample properties of the estimator have been studied by computer simulations. The model and associated methods of parametric inference have been applied to the analysis of proliferation and differentiation of cultured O-2A progenitor cells that play a key role in the development of the central nervous system. It follows from this analysis that the time to division of the progenitor cell and the time to its differentiation (into an oligodendrocyte) are not identically distributed. This biological finding suggests that a molecular event determining the type of cell transformation is more likely to occur at the start rather than at the end of the mitotic cycle.
BACKGROUND: To identify differentially expressed genes, it is standard practice to test a two-sample hypothesis for each gene with a proper adjustment for multiple testing. Such tests are essentially univariate and disregard the multidimensional structure of microarray data. A more general two-sample hypothesis is formulated in terms of the joint distribution of any sub-vector of expression signals. RESULTS: By building on an earlier proposed multivariate test statistic, we propose a new algorithm for identifying differentially expressed gene combinations. The algorithm includes an improved random search procedure designed to generate candidate gene combinations of a given size. Cross-validation is used to provide replication stability of the search procedure. A permutation two-sample test is used for significance testing. We design a multiple testing procedure to control the family-wise error rate (FWER) when selecting significant combinations of genes that result from a successive selection procedure. A target set of genes is composed of all significant combinations selected via random search. CONCLUSIONS: A new algorithm has been developed to identify differentially expressed gene combinations. The performance of the proposed search-and-testing procedure has been evaluated by computer simulations and analysis of replicated Affymetrix gene array data on age-related changes in gene expression in the inner ear of CBA mice.
This paper considers the utility of statistical goodness of fit testing in the context of mechanistic models of carcinogenesis. Two stochastic models of carcinogenesis were tested with several sets of experimental and epidemiological data using a formal goodness of fit test specially designed to accommodate censored observations: these were the two-stage model allowing for clonal expansion of initiated cells and its simpler version with gamma distributed promotion time. The results of this application, supplemented by visual examination of local likelihood kernel estimates of the hazard function and the corresponding model-based estimates, show that mechanistic models of carcinogenesis provide a good fit to the data in the majority of cases under study.