Search PubMed⌕ Search

Biomedical subjects

G Stolovitzky

Publications and source records attributed to G Stolovitzky.

At least 19 recordsLinked to original sources

Quantitative noise analysis for gene expression microarray experiments.

A major challenge in DNA microarray analysis is to effectively dissociate actual gene expression values from experimental noise. We report here a detailed noise analysis for oligonuleotide-based microarray experiments involving reverse transcription, generation of labeled cRNA (target) through in vitro transcription, and hybridization of the target to the probe immobilized on the substrate. By designing sets of replicate experiments that bifurcate at different steps of the assay, we are able to separate the noise caused by sample preparation and the hybridization processes. We quantitatively characterize the strength of these different sources of noise and their respective dependence on the gene expression level. We find that the sample preparation noise is small, implying that the amplification process during the sample preparation is relatively accurate. The hybridization noise is found to have very strong dependence on the expression level, with different characteristics for the low and high expression values. The hybridization noise characteristics at the high expression regime are mostly Poisson-like, whereas its characteristics for the small expression levels are more complex, probably due to cross-hybridization. A method to evaluate the significance of gene expression fold changes based on noise characteristics is proposed.

Gene Expression↗

Catalytic tempering: A method for sampling rough energy landscapes by Monte Carlo.

A new Monte Carlo algorithm is presented for the efficient sampling of the Boltzmann distribution of configurations of systems with rough energy landscapes. The method is based on the introduction of a fictitious coordinate y so that the dimensionality of the system is increased by one. This augmented system has a potential surface and a temperature that is made to depend on the new coordinate y in such a way that for a small strip of the y space, called the "normal region," the temperature is set equal to the temperature desired and the potential is the original rough energy potential. To enhance barrier crossing outside the "normal region," the energy barriers are reduced by truncation (with preservation of the potential minima) and the temperature is made to increase with ||y ||. The method, called catalytic tempering or CAT, is found to greatly improve the rate of convergence of Monte Carlo sampling in model systems and to eliminate the quasi-ergodic behavior often found in the sampling of rough energy landscapes.

Journal Article↗

Systematic and fully automated identification of protein sequence patterns.

We present an efficient algorithm to systematically and automatically identify patterns in protein sequence families. The procedure is based on the Splash deterministic pattern discovery algorithm and on a framework to assess the statistical significance of patterns. We demonstrate its application to the fully automated discovery of patterns in 974 PROSITE families (the complete subset of PROSITE families which are defined by patterns and contain DR records). Splash generates patterns with better specificity and undiminished sensitivity, or vice versa, in 28% of the families; identical statistics were obtained in 48% of the families, worse statistics in 15%, and mixed behavior in the remaining 9%. In about 75% of the cases, Splash patterns identify sequence sites that overlap more than 50% with the corresponding PROSITE pattern. The procedure is sufficiently rapid to enable its use for daily curation of existing motif and profile databases. Third, our results show that the statistical significance of discovered patterns correlates well with their biological significance. The trypsin subfamily of serine proteases is used to illustrate this method's ability to exhaustively discover all motifs in a family that are statistically and biologically significant. Finally, we discuss applications of sequence patterns to multiple sequence alignment and the training of more sensitive score-based motif models, akin to the procedure used by PSI-BLAST. All results are available at httpl//www.research.ibm.com/spat/.

Algorithms↗

Analysis of gene expression microarrays for phenotype classification.

Several microarray technologies that monitor the level of expression of a large number of genes have recently emerged. Given DNA-microarray data for a set of cells characterized by a given phenotype and for a set of control cells, an important problem is to identify "patterns" of gene expression that can be used to predict cell phenotype. The potential number of such patterns is exponential in the number of genes. In this paper, we propose a solution to this problem based on a supervised learning algorithm, which differs substantially from previous schemes. It couples a complex, non-linear similarity metric, which maximizes the probability of discovering discriminative gene expression patterns, and a pattern discovery algorithm called SPLASH. The latter discovers efficiently and deterministically all statistically significant gene expression patterns in the phenotype set. Statistical significance is evaluated based on the probability of a pattern to occur by chance in the control set. Finally, a greedy set covering algorithm is used to select an optimal subset of statistically significant patterns, which form the basis for a standard likelihood ratio classification scheme. We analyze data from 60 human cancer cell lines using this method, and compare our results with those of other supervised learning schemes. Different phenotypes are studied. These include cancer morphologies (such as melanoma), molecular targets (such as mutations in the p53 gene), and therapeutic targets related to the sensitivity to an anticancer compounds. We also analyze a synthetic data set that shows that this technique is especially well suited for the analysis of sub-phenotype mixtures. For complex phenotypes, such as p53, our method produces an encouragingly low rate of false positives and false negatives and seems to outperform the others. Similar low rates are reported when predicting the efficacy of experimental anticancer compounds. This counts among the first reported studies where drug efficacy has been successfully predicted from large-scale expression data analysis.

Algorithms↗

Compositional heterogeneity within, and uniformity between, DNA sequences of yeast chromosomes.

The heterogeneity within, and similarities between, yeast chromosomes are studied. For the former, we show by the size distribution of domains, coding density, size distribution of open reading frames, spatial power spectra, and deviation from binomial distribution for C + G% in large moving windows that there is a strong deviation of the yeast sequences from random sequences. For the latter, not only do we graphically illustrate the similarity for the above mentioned statistics, but we also carry out a rigorous analysis of variance (ANOVA) test. The hypothesis that all yeast chromosomes are similar cannot be rejected by this test. We examine the two possible explanations of this interchromosomal uniformity: a common origin, such as genome-wide duplication (polyploidization), and a concerted evolutionary process.

Analysis of Variance↗

Mnemonics for variability: remembering food delay.

Three experiments with White Carneaux pigeons (Columba livia) investigated memory and decision processes under fixed and variable reinforcement intervals. Response rate was measured during the unreinforced trials in the discrete-trial peak procedure in which reinforced trials were mixed with long unreinforced trials. Two decision models differing in assumptions about memory constraints are reviewed. In the complete-memory model (J. Gibbon, R.M. Church, S. Fairhurst, & A. Kacelnik, 1988), all interreinforcement intervals were remembered, whereas in the minimax model (D. Brunner, A. Kacelnik, & J. Gibbon, 1996), only estimates of the shortest and longest possible reinforcement times were remembered. Both models accommodated some features of response rate as a function of trial time, but only the second was compatible with the observed cessation of responding.

Animals↗

Efficiency of DNA replication in the polymerase chain reaction.

A detailed quantitative kinetic model for the polymerase chain reaction (PCR) is developed, which allows us to predict the probability of replication of a DNA molecule in terms of the physical parameters involved in the system. The important issue of the determination of the number of PCR cycles during which this probability can be considered to be a constant is solved within the framework of the model. New phenomena of multimodality and scaling behavior in the distribution of the number of molecules after a given number of PCR cycles are presented. The relevance of the model for quantitative PCR is discussed, and a novel quantitative PCR technique is proposed.

DNA Replication↗