Search PubMed⌕ Search

Biomedical subjects

Sandrine Dudoit

Publications and source records attributed to Sandrine Dudoit.

9 recordsLinked to original sources

transfactor: transcription factor activity estimation via probabilistic gene expression deconvolution.

Gene expression is a primary modality being studied to differentiate between biological cells. Contemporary single-cell studies simultaneously measure genome-wide transcription levels for thousands of individual cells in a single experiment. While the characterization of cell population differences has often occurred through differential gene expression analysis, tiny effect sizes become statistically significant when thousands of cells are available for each population, compromising biological interpretation. Moreover, these large studies have spurred the development of methods to infer gene regulatory networks (GRNs) directly from the data, and GRN databases are becoming more comprehensive. In this work, we propose a statistical model for gene expression measures and an inference method that leverage GRNs to deconvolve transcription factor (TF) activity from gene expression, by probabilistically assigning mRNA molecules to TFs. This shifts the paradigm from investigating gene expression differences to regulatory differences at the level of TF activity, aiding interpretation and allowing prioritization of a limited number of TFs responsible for significant contributions to the observed gene expression differences. The inferred TF activities result in intuitive prioritization of TFs in terms of the (difference in) estimated number of molecules they produce, in contrast to other widely used methods relying on arbitrary enrichment scores. Our model allows the incorporation of prior information on the regulatory potential between each TF and target gene and is able to deal with both repressing and activating interactions. We compare our approach to other TF activity estimation methods using two simulation experiments and two case studies. Single-cell RNA-sequencing; TF activity; bioinformatics; GRN.

Transcription Factors↗

A latent activated olfactory stem cell state revealed by single-cell transcriptomic and epigenomic profiling.

The olfactory epithelium is one of the few regions of the nervous system that sustains neurogenesis throughout life. Its experimental accessibility makes it especially tractable for studying molecular mechanisms that drive neural regeneration in response to injury. In this study, we used single-cell sequencing to identify transcriptional and epigenetic processes involved in determining olfactory epithelial stem cell fate during injury-induced regeneration. By combining gene expression and accessible chromatin profiles of individual lineage-traced olfactory stem cells, we identified transcriptional heterogeneity among activated stem cells at a stage when cell fates are being specified. We further identified a subset of resting cells that appears poised for activation, characterized by accessible chromatin around silent genes prior to their expression in response to injury. These results provide evidence for a latent activated stem cell state in which a subset of quiescent olfactory epithelial stem cells are epigenetically primed to support injury-induced regeneration.

Animals↗

Bagging to improve the accuracy of a clustering procedure.

MOTIVATION: The microarray technology is increasingly being applied in biological and medical research to address a wide range of problems such as the classification of tumors. An important statistical question associated with tumor classification is the identification of new tumor classes using gene expression profiles. Essential aspects of this clustering problem include identifying accurate partitions of the tumor samples into clusters and assessing the confidence of cluster assignments for individual samples. RESULTS: Two new resampling methods, inspired from bagging in prediction, are proposed to improve and assess the accuracy of a given clustering procedure. In these ensemble methods, a partitioning clustering procedure is applied to bootstrap learning sets and the resulting multiple partitions are combined by voting or the creation of a new dissimilarity matrix. As in prediction, the motivation behind bagging is to reduce variability in the partitioning results via averaging. The performances of the new and existing methods were compared using simulated data and gene expression data from two recently published cancer microarray studies. The bagged clustering procedures were in general at least as accurate and often substantially more accurate than a single application of the partitioning clustering procedure. A valuable by-product of bagged clustering are the cluster votes which can be used to assess the confidence of cluster assignments for individual observations. SUPPLEMENTARY INFORMATION: For supplementary information on datasets, analyses, and software, consult http://www.stat.berkeley.edu/~sandrine and http://www.bioconductor.org.

Algorithms↗

Open source software for the analysis of microarray data.

DNA microarray assays represent the first widely used application that attempts to build upon the information provided by genome projects in the study of biological questions. One of the greatest challenges with working with microarrays is collecting, managing, and analyzing data. Although several commercial and noncommercial solutions exist, there is a growing body of freely available, open source software that allows users to analyze data using a host of existing techniques and to develop their own and integrate them within the system. Here we review three of the most widely used and comprehensive systems, the statistical analysis tools written in R through the Bioconductor project (http://www.bioconductor.org), the Java-based TM4 software system available from The Institute for Genomic Research (http://www.tigr.org/software), and BASE, the Web-based system developed at Lund University (http://base.thep.lu.se).

Database Management Systems↗

Diversity, topographic differentiation, and positional memory in human fibroblasts.

A fundamental feature of the architecture and functional design of vertebrate animals is a stroma, composed of extracellular matrix and mesenchymal cells, which provides a structural scaffold and conduit for blood and lymphatic vessels, nerves, and leukocytes. Reciprocal interactions between mesenchymal and epithelial cells are known to play a critical role in orchestrating the development and morphogenesis of tissues and organs, but the roles played by specific stromal cells in controlling the design and function of tissues remain poorly understood. The principal cells of stromal tissue are called fibroblasts, a catch-all designation that belies their diversity. We characterized genome-wide patterns of gene expression in cultured fetal and adult human fibroblasts derived from skin at different anatomical sites. Fibroblasts from each site displayed distinct and characteristic transcriptional patterns, suggesting that fibroblasts at different locations in the body should be considered distinct differentiated cell types. Notable groups of differentially expressed genes included some implicated in extracellular matrix synthesis, lipid metabolism, and cell signaling pathways that control proliferation, cell migration, and fate determination. Several genes implicated in genetic diseases were found to be expressed in fibroblasts in an anatomic pattern that paralleled the phenotypic defects. Finally, adult fibroblasts maintained key features of HOX gene expression patterns established during embryogenesis, suggesting that HOX genes may direct topographic differentiation and underlie the detailed positional memory in fibroblasts.

Cell Differentiation↗

A prediction-based resampling method for estimating the number of clusters in a dataset.

BACKGROUND: Microarray technology is increasingly being applied in biological and medical research to address a wide range of problems, such as the classification of tumors. An important statistical problem associated with tumor classification is the identification of new tumor classes using gene-expression profiles. Two essential aspects of this clustering problem are: to estimate the number of clusters, if any, in a dataset; and to allocate tumor samples to these clusters, and assess the confidence of cluster assignments for individual samples. Here we address the first of these problems. RESULTS: We have developed a new prediction-based resampling method, Clest, to estimate the number of clusters in a dataset. The performance of the new and existing methods were compared using simulated data and gene-expression data from four recently published cancer microarray studies. Clest was generally found to be more accurate and robust than the six existing methods considered in the study. CONCLUSIONS: Focusing on prediction accuracy in conjunction with resampling produces accurate and robust estimates of the number of clusters.

Algorithms↗

Normalization for cDNA microarray data: a robust composite method addressing single and multiple slide systematic variation.

There are many sources of systematic variation in cDNA microarray experiments which affect the measured gene expression levels (e.g. differences in labeling efficiency between the two fluorescent dyes). The term normalization refers to the process of removing such variation. A constant adjustment is often used to force the distribution of the intensity log ratios to have a median of zero for each slide. However, such global normalization approaches are not adequate in situations where dye biases can depend on spot overall intensity and/or spatial location within the array. This article proposes normalization methods that are based on robust local regression and account for intensity and spatial dependence in dye biases for different types of cDNA microarray experiments. The selection of appropriate controls for normalization is discussed and a novel set of controls (microarray sample pool, MSP) is introduced to aid in intensity-dependent normalization. Lastly, to allow for comparisons of expression levels across slides, a robust method based on maximum likelihood estimation is proposed to adjust for scale differences among slides.

Animals↗

Stereotyped and specific gene expression programs in human innate immune responses to bacteria.

The innate immune response is crucial for defense against microbial pathogens. To investigate the molecular choreography of this response, we carried out a systematic examination of the gene expression program in human peripheral blood mononuclear cells responding to bacteria and bacterial products. We found a remarkably stereotyped program of gene expression induced by bacterial lipopolysaccharide and diverse killed bacteria. An intricately choreographed expression program devoted to communication between cells was a prominent feature of the response. Other features suggested a molecular program for commitment of antigen-presenting cells to antigens captured in the context of bacterial infection. Despite the striking similarities, there were qualitative and quantitative differences in the responses to different bacteria. Modulation of this host-response program by bacterial virulence mechanisms was an important source of variation in the response to different bacteria.

Bacteria↗

Gene expression patterns in human liver cancers.

Hepatocellular carcinoma (HCC) is a leading cause of death worldwide. Using cDNA microarrays to characterize patterns of gene expression in HCC, we found consistent differences between the expression patterns in HCC compared with those seen in nontumor liver tissues. The expression patterns in HCC were also readily distinguished from those associated with tumors metastatic to liver. The global gene expression patterns intrinsic to each tumor were sufficiently distinctive that multiple tumor nodules from the same patient could usually be recognized and distinguished from all the others in the large sample set on the basis of their gene expression patterns alone. The distinctive gene expression patterns are characteristic of the tumors and not the patient; the expression programs seen in clonally independent tumor nodules in the same patient were no more similar than those in tumors from different patients. Moreover, clonally related tumor masses that showed distinct expression profiles were also distinguished by genotypic differences. Some features of the gene expression patterns were associated with specific phenotypic and genotypic characteristics of the tumors, including growth rate, vascular invasion, and p53 overexpression.

Carcinoma, Hepatocellular↗