Search PubMed⌕ Search

Biomedical subjects

Olga G Troyanskaya

Publications and source records attributed to Olga G Troyanskaya.

At least 19 recordsLinked to original sources

Single-cell profiling reveals epithelial and immune responses in BK polyomavirus-infected human kidney biopsies.

INTRODUCTIONBK polyomavirus (BKV) infection is associated with injury and subsequent graft loss due to the extent of injury or rejection. However, the molecular mechanisms driving injury and subsequent adverse outcomes remain poorly understood.METHODSIn a cross-sectional study, single-cell RNA-seq from kidney allograft biopsies was used to assess cell type-specific responses between uninfected controls and 2 distinct phases of BKV infection: peaking (increasing viral blood titers) and resolving (decreasing viral titers following immunosuppression reduction).RESULTSGenes upregulated in BK viral nephropathy (BKVN) were enriched for polyomavirus infection hallmarks, including ribosome biogenesis, translation, and energy restructuring. Additionally, enriched pathways included wound healing, cellular stress, antigen presentation and immune signaling. Even without BKVN (peaking BK viremia alone), epithelial cells expressed signatures for wound healing, cellular stress, and extracellular matrix remodeling. In vivo tubular cell responses at single-cell resolution were validated against single cell transcriptomic data of BKV-infected cells in a cell culture model. Despite similarities, in vivo tubular cells underwent metabolic adaptation favoring fatty acid oxidation and proinflammatory responses not observed in culture models, likely due to an absent innate and adaptive immune system. Despite lymphopenia and immunosuppressive therapies, the proportion of recipient-derived intrarenal adaptive immune cells was increased in biopsies associated with peaking viremia alongside activation of innate immune responses. Adaptive immune cells exhibited persistent inflammatory signaling and remodeling of energy metabolism during the resolving phase of infection.CONCLUSIONThese not previously reported insights into BKV-associated injury may have implications for clinical management and improved allograft outcomes.

Humans↗

Innate immune molecular landscape following controlled human influenza virus infection.

Viral infections can induce prolonged changes in innate immunity. Here, we use blood samples from a human influenza H3N2 challenge study (NCT03883113) to perform comprehensive multi-omics analyses. We detect remodeling of immune programs in circulating innate immune cells that persist after resolution of the infection. We find changes associated with suppressed inflammation, including decreased cytokine and AP-1 gene expression as well as decreased accessibility at AP-1 targets and interleukin-related gene promoter regions. We also find decreased histone deacetylase gene expression, increased MAP kinase gene expression, and increased accessibility at interferon-related gene promoter regions. Genes involved in inflammation and methylation remodeling show modulation of gene-chromatin site regulatory circuit activity. These results reveal a coordinated rewiring of the molecular landscape in innate immune cells induced by mild influenza virus infection.

Humans↗

3D chromatin structures precede genome activation in Drosophila embryogenesis.

3D chromatin structure is critical for the regulation of gene expression during development. Here we used Micro-C assays at 100-bp resolution to map genome organization in Drosophila melanogaster throughout the first half of embryogenesis. These high-resolution contact maps reveal fine-scale features such as loops and boundaries delineating topologically associating domains. Notably, we observe that 3D chromatin structures form prior to zygotic genome activation and persist during successive mitotic cycles. Integrative analysis with 149 public chromatin immunoprecipitation sequencing (ChIP-seq) datasets identifies four classes of chromatin structuring elements, including a distinct group enriched for GAGA-associated factor (GAF) and Zelda binding, associated with developmental-gene regulation. These elements are mitotically retained and exhibit sequence and structure similarity between D. melanogaster and D. virilis. We propose that 3D chromatin organization in the pre-cellular embryo facilitates deployment of developmentally regulated genes during Drosophila embryogenesis.

Animals↗

Modeling complex genetic interactions in a simple eukaryotic genome: actin displays a rich spectrum of complex haploinsufficiencies.

Multigenic influences are major contributors to human genetic disorders. Since humans are highly polymorphic, there are a high number of possible detrimental, multiallelic gene pairs. The actin cytoskeleton of yeast was used to determine the potential for deleterious bigenic interactions; approximately 4800 complex hemizygote strains were constructed between an actin-null allele and the nonessential gene deletion collection. We found 208 genes that have deleterious complex haploinsufficient (CHI) interactions with actin. This set is enriched for genes with gene ontology terms shared with actin, including several actin-binding protein genes, and nearly half of the CHI genes have defects in actin organization when deleted. Interactions were frequently seen with genes for multiple components of a complex or with genes involved in the same function. For example, many of the genes for the large ribosomal subunit (RPLs) were CHI with act1Delta and had actin organization defects when deleted. This was generally true of only one RPL paralog of apparently duplicate genes, suggesting functional specialization between ribosomal genes. In many cases, CHI interactions could be attributed to localized defects on the actin protein. Spatial congruence in these data suggest that the loss of binding to specific actin-binding proteins causes subsets of CHI interactions.

Actins↗

Functional analysis of gene duplications in Saccharomyces cerevisiae.

Gene duplication can occur on two scales: whole-genome duplications (WGD) and smaller-scale duplications (SSD) involving individual genes or genomic segments. Duplication may result in functionally redundant genes or diverge in function through neofunctionalization or subfunctionalization. The effect of duplication scale on functional evolution has not yet been explored, probably due to the lack of global knowledge of protein function and different times of duplication events. To address this question, we used integrated Bayesian analysis of diverse functional genomic data to accurately evaluate the extent of functional similarity and divergence between paralogs on a global scale. We found that paralogs resulting from the whole-genome duplication are more likely to share interaction partners and biological functions than smaller-scale duplicates, independent of sequence similarity. In addition, WGD paralogs show lower frequency of essential genes and higher synthetic lethality rate, but instead diverge more in expression pattern and upstream regulatory region. Thus, our analysis demonstrates that WGD paralogs generally have similar compensatory functions but diverging expression patterns, suggesting a potential of distinct evolutionary scenarios for paralogs that arose through different duplication mechanisms. Furthermore, by identifying these functional disparities between the two types of duplicates, we reconcile previous disputes on the relationship between sequence divergence and expression divergence or essentiality.

Bayes Theorem↗

GOLEM: an interactive graph-based gene-ontology navigation and analysis tool.

BACKGROUND: The Gene Ontology has become an extremely useful tool for the analysis of genomic data and structuring of biological knowledge. Several excellent software tools for navigating the gene ontology have been developed. However, no existing system provides an interactively expandable graph-based view of the gene ontology hierarchy. Furthermore, most existing tools are web-based or require an Internet connection, will not load local annotations files, and provide either analysis or visualization functionality, but not both. RESULTS: To address the above limitations, we have developed GOLEM (Gene Ontology Local Exploration Map), a visualization and analysis tool for focused exploration of the gene ontology graph. GOLEM allows the user to dynamically expand and focus the local graph structure of the gene ontology hierarchy in the neighborhood of any chosen term. It also supports rapid analysis of an input list of genes to find enriched gene ontology terms. The GOLEM application permits the user either to utilize local gene ontology and annotations files in the absence of an Internet connection, or to access the most recent ontology and annotation information from the gene ontology webpage. GOLEM supports global and organism-specific searches by gene ontology term name, gene ontology id and gene name. CONCLUSION: GOLEM is a useful software tool for biologists interested in visualizing the local directed acyclic graph structure of the gene ontology hierarchy and searching for gene ontology terms enriched in genes of interest. It is freely available both as an application and as an applet at http://function.princeton.edu/GOLEM.

Computer Graphics↗

A scalable method for integration and functional analysis of multiple microarray datasets.

MOTIVATION: The diverse microarray datasets that have become available over the past several years represent a rich opportunity and challenge for biological data mining. Many supervised and unsupervised methods have been developed for the analysis of individual microarray datasets. However, integrated analysis of multiple datasets can provide a broader insight into genetic regulation of specific biological pathways under a variety of conditions. RESULTS: To aid in the analysis of such large compendia of microarray experiments, we present Microarray Experiment Functional Integration Technology (MEFIT), a scalable Bayesian framework for predicting functional relationships from integrated microarray datasets. Furthermore, MEFIT predicts these functional relationships within the context of specific biological processes. All results are provided in the context of one or more specific biological functions, which can be provided by a biologist or drawn automatically from catalogs such as the Gene Ontology (GO). Using MEFIT, we integrated 40 Saccharomyces cerevisiae microarray datasets spanning 712 unique conditions. In tests based on 110 biological functions drawn from the GO biological process ontology, MEFIT provided a 5% or greater performance increase for 54 functions, with a 5% or more decrease in performance in only two functions.

Algorithms↗

Finding function: evaluation methods for functional genomic data.

BACKGROUND: Accurate evaluation of the quality of genomic or proteomic data and computational methods is vital to our ability to use them for formulating novel biological hypotheses and directing further experiments. There is currently no standard approach to evaluation in functional genomics. Our analysis of existing approaches shows that they are inconsistent and contain substantial functional biases that render the resulting evaluations misleading both quantitatively and qualitatively. These problems make it essentially impossible to compare computational methods or large-scale experimental datasets and also result in conclusions that generalize poorly in most biological applications. RESULTS: We reveal issues with current evaluation methods here and suggest new approaches to evaluation that facilitate accurate and representative characterization of genomic methods and data. Specifically, we describe a functional genomics gold standard based on curation by expert biologists and demonstrate its use as an effective means of evaluation of genomic approaches. Our evaluation framework and gold standard are freely available to the community through our website. CONCLUSION: Proper methods for evaluating genomic data and computational approaches will determine how much we, as a community, are able to learn from the wealth of available data. We propose one possible solution to this problem here but emphasize that this topic warrants broader community discussion.

Algorithms↗

Global analysis of gene function in yeast by quantitative phenotypic profiling.

We present a method for the global analysis of the function of genes in budding yeast based on hierarchical clustering of the quantitative sensitivity profiles of the 4756 strains with individual homozygous deletion of nonessential genes to a broad range of cytotoxic or cytostatic agents. This method is superior to other global methods of identifying the function of genes involved in the various DNA repair and damage checkpoint pathways as well as other interrogated functions. Analysis of the phenotypic profiles of the 51 diverse treatments places a total of 860 genes of unknown function in clusters with genes of known function. We demonstrate that this can not only identify the function of unknown genes but can also suggest the mechanism of action of the agents used. This method will be useful when used alone and in conjunction with other global approaches to identify gene function in yeast.

Cluster Analysis↗

Hierarchical multi-label prediction of gene function.

MOTIVATION: Assigning functions for unknown genes based on diverse large-scale data is a key task in functional genomics. Previous work on gene function prediction has addressed this problem using independent classifiers for each function. However, such an approach ignores the structure of functional class taxonomies, such as the Gene Ontology (GO). Over a hierarchy of functional classes, a group of independent classifiers where each one predicts gene membership to a particular class can produce a hierarchically inconsistent set of predictions, where for a given gene a specific class may be predicted positive while its inclusive parent class is predicted negative. Taking the hierarchical structure into account resolves such inconsistencies and provides an opportunity for leveraging all classifiers in the hierarchy to achieve higher specificity of predictions. RESULTS: We developed a Bayesian framework for combining multiple classifiers based on the functional taxonomy constraints. Using a hierarchy of support vector machine (SVM) classifiers trained on multiple data types, we combined predictions in our Bayesian framework to obtain the most probable consistent set of predictions. Experiments show that over a 105-node subhierarchy of the GO, our Bayesian framework improves predictions for 93 nodes. As an additional benefit, our method also provides implicit calibration of SVM margin outputs to probabilities. Using this method, we make function predictions for multiple proteins, and experimentally confirm predictions for proteins involved in mitosis. SUPPLEMENTARY INFORMATION: Results for the 105 selected GO classes and predictions for 1059 unknown genes are available at: http://function.princeton.edu/genesite/ CONTACT: ogt@cs.princeton.edu.

Algorithms↗

Discovery of biological networks from diverse functional genomic data.

We have developed a general probabilistic system for query-based discovery of pathway-specific networks through integration of diverse genome-wide data. This framework was validated by accurately recovering known networks for 31 biological processes in Saccharomyces cerevisiae and experimentally verifying predictions for the process of chromosomal segregation. Our system, bioPIXIE, a public, comprehensive system for integration, analysis, and visualization of biological network predictions for S. cerevisiae, is freely accessible over the worldwide web.

Bayes Theorem↗

Visualization-based discovery and analysis of genomic aberrations in microarray data.

BACKGROUND: Chromosomal copy number changes (aneuploidies) play a key role in cancer progression and molecular evolution. These copy number changes can be studied using microarray-based comparative genomic hybridization (array CGH) or gene expression microarrays. However, accurate identification of amplified or deleted regions requires a combination of visual and computational analysis of these microarray data. RESULTS: We have developed ChARMView, a visualization and analysis system for guided discovery of chromosomal abnormalities from microarray data. Our system facilitates manual or automated discovery of aneuploidies through dynamic visualization and integrated statistical analysis. ChARMView can be used with array CGH and gene expression microarray data, and multiple experiments can be viewed and analyzed simultaneously. CONCLUSION: ChARMView is an effective and accurate visualization and analysis system for recognizing even small aneuploidies or subtle expression biases, identifying recurring aberrations in sets of experiments, and pinpointing functionally relevant copy number changes. ChARMView is freely available under the GNU GPL at http://function.princeton.edu/ChARMView.

Algorithms↗

Visualization methods for statistical analysis of microarray clusters.

BACKGROUND: The most common method of identifying groups of functionally related genes in microarray data is to apply a clustering algorithm. However, it is impossible to determine which clustering algorithm is most appropriate to apply, and it is difficult to verify the results of any algorithm due to the lack of a gold-standard. Appropriate data visualization tools can aid this analysis process, but existing visualization methods do not specifically address this issue. RESULTS: We present several visualization techniques that incorporate meaningful statistics that are noise-robust for the purpose of analyzing the results of clustering algorithms on microarray data. This includes a rank-based visualization method that is more robust to noise, a difference display method to aid assessments of cluster quality and detection of outliers, and a projection of high dimensional data into a three dimensional space in order to examine relationships between clusters. Our methods are interactive and are dynamically linked together for comprehensive analysis. Further, our approach applies to both protein and gene expression microarrays, and our architecture is scalable for use on both desktop/laptop screens and large-scale display devices. This methodology is implemented in GeneVAnD (Genomic Visual ANalysis of Datasets) and is available at http://function.princeton.edu/GeneVAnD. CONCLUSION: Incorporating relevant statistical information into data visualizations is key for analysis of large biological datasets, particularly because of high levels of noise and the lack of a gold-standard for comparisons. We developed several new visualization techniques and demonstrated their effectiveness for evaluating cluster quality and relationships between clusters.

Algorithms↗

Putting microarrays in a context: integrated analysis of diverse biological data.

In recent years, multiple types of high-throughput functional genomic data that facilitate rapid functional annotation of sequenced genomes have become available. Gene expression microarrays are the most commonly available source of such data. However, genomic data often sacrifice specificity for scale, yielding very large quantities of relatively lower-quality data than traditional experimental methods. Thus sophisticated analysis methods are necessary to make accurate functional interpretation of these large-scale data sets. This review presents an overview of recently developed methods that integrate the analysis of microarray data with sequence, interaction, localisation and literature data, and further outlines current challenges in the field. The focus of this review is on the use of such methods for gene function prediction, understanding of protein regulation and modelling of biological networks.

Algorithms↗

Accurate detection of aneuploidies in array CGH and gene expression microarray data.

MOTIVATION: Chromosomal copy number changes (aneuploidies) are common in cell populations that undergo multiple cell divisions including yeast strains, cell lines and tumor cells. Identification of aneuploidies is critical in evolutionary studies, where changes in copy number serve an adaptive purpose, as well as in cancer studies, where amplifications and deletions of chromosomal regions have been identified as a major pathogenetic mechanism. Aneuploidies can be studied on whole-genome level using array CGH (a microarray-based method that measures the DNA content), but their presence also affects gene expression. In gene expression microarray analysis, identification of copy number changes is especially important in preventing aberrant biological conclusions based on spurious gene expression correlation or masked phenotypes that arise due to aneuploidies. Previously suggested approaches for aneuploidy detection from microarray data mostly focus on array CGH, address only whole-chromosome or whole-arm copy number changes, and rely on thresholds or other heuristics, making them unsuitable for fully automated general application to gene expression datasets. There is a need for a general and robust method for identification of aneuploidies of any size from both array CGH and gene expression microarray data. RESULTS: We present ChARM (Chromosomal Aberration Region Miner), a robust and accurate expectation-maximization based method for identification of segmental aneuploidies (partial chromosome changes) from gene expression and array CGH microarray data. Systematic evaluation of the algorithm on synthetic and biological data shows that the method is robust to noise, aneuploidal segment size and P-value cutoff. Using our approach, we identify known chromosomal changes and predict novel potential segmental aneuploidies in commonly used yeast deletion strains and in breast cancer. ChARM can be routinely used to identify aneuploidies in array CGH datasets and to screen gene expression data for aneuploidies or array biases. Our methodology is sensitive enough to detect statistically significant and biologically relevant aneuploidies even when expression or DNA content changes are subtle as in mixed populations of cells. AVAILABILITY: Code available by request from the authors and on Web supplement at http://function.cs.princeton.edu/ChARM/

Algorithms↗

Systemic and cell type-specific gene expression patterns in scleroderma skin.

We used DNA microarrays representing >12,000 human genes to characterize gene expression patterns in skin biopsies from individuals with a diagnosis of systemic sclerosis with diffuse scleroderma. We found consistent differences in the patterns of gene expression between skin biopsies from individuals with scleroderma and those from normal, unaffected individuals. The biopsies from affected individuals showed nearly indistinguishable patterns of gene expression in clinically affected and clinically unaffected tissue, even though these were clearly distinguishable from the patterns found in similar tissue from unaffected individuals. Genes characteristically expressed in endothelial cells, B lymphocytes, and fibroblasts showed differential expression between scleroderma and normal biopsies. Analysis of lymphocyte populations in scleroderma skin biopsies by immunohistochemistry suggest the B lymphocyte signature observed on our arrays is from CD20+ B cells. These results provide evidence that scleroderma has systemic manifestations that affect multiple cell types and suggests genes that could be used as potential markers for the disease.

Adult↗

Endothelial cell diversity revealed by global expression profiling.

The vascular system is locally specialized to accommodate widely varying blood flow and pressure and the distinct needs of individual tissues. The endothelial cells (ECs) that line the lumens of blood and lymphatic vessels play an integral role in the regional specialization of vascular structure and physiology. However, our understanding of EC diversity is limited. To explore EC specialization on a global scale, we used DNA microarrays to determine the expression profile of 53 cultured ECs. We found that ECs from different blood vessels and microvascular ECs from different tissues have distinct and characteristic gene expression profiles. Pervasive differences in gene expression patterns distinguish the ECs of large vessels from microvascular ECs. We identified groups of genes characteristic of arterial and venous endothelium. Hey2, the human homologue of the zebrafish gene gridlock, was selectively expressed in arterial ECs and induced the expression of several arterial-specific genes. Several genes critical in the establishment of left/right asymmetry were expressed preferentially in venous ECs, suggesting coordination between vascular differentiation and body plan development. Tissue-specific expression patterns in different tissue microvascular ECs suggest they are distinct differentiated cell types that play roles in the local physiology of their respective organs and tissues.

Cells, Cultured↗