Search PubMed⌕ Search

Biomedical subjects

Soumya Raychaudhuri

Publications and source records attributed to Soumya Raychaudhuri.

7 recordsLinked to original sources

Investigating hypoxic tumor physiology through gene expression patterns.

Clinical evidence shows that tumor hypoxia is an independent prognostic indicator of poor patient outcome. Hypoxic tumors have altered physiologic processes, including increased regions of angiogenesis, increased local invasion, increased distant metastasis and altered apoptotic programs. Since hypoxia is a potent controller of gene expression, identifying hypoxia-regulated genes is a means to investigate the molecular response to hypoxic stress. Traditional experimental approaches have identified physiologic changes in hypoxic cells. Recent studies have identified hypoxia-responsive genes that may define the mechanism(s) underlying these physiologic changes. For example, the regulation of glycolytic genes by hypoxia can explain some characteristics of the Warburg effect. The converse of this logic is also true. By identifying new classes of hypoxia-regulated gene(s), we can infer the physiologic pressures that require the induction of these genes and their protein products. Furthermore, these physiologically driven hypoxic gene expression changes give us insight as to the poor outcome of patients with hypoxic tumors. Approximately 1-1.5% of the genome is transcriptionally responsive to hypoxia. However, there is significant heterogeneity in the transcriptional response to hypoxia between different cell types. Moreover, the coordinated change in the expression of families of genes supports the model of physiologic pressure leading to expression changes. Understanding the evolutionary pressure to develop a 'hypoxic response' provides a framework to investigate the biology of the hypoxic tumor microenvironment.

Animals↗

The computational analysis of scientific literature to define and recognize gene expression clusters.

A limitation of many gene expression analytic approaches is that they do not incorporate comprehensive background knowledge about the genes into the analysis. We present a computational method that leverages the peer-reviewed literature in the automatic analysis of gene expression data sets. Including the literature in the analysis of gene expression data offers an opportunity to incorporate functional information about the genes when defining expression clusters. We have created a method that associates gene expression profiles with known biological functions. Our method has two steps. First, we apply hierarchical clustering to the given gene expression data set. Secondly, we use text from abstracts about genes to (i) resolve hierarchical cluster boundaries to optimize the functional coherence of the clusters and (ii) recognize those clusters that are most functionally coherent. In the case where a gene has not been investigated and therefore lacks primary literature, articles about well-studied homologous genes are added as references. We apply our method to two large gene expression data sets with different properties. The first contains measurements for a subset of well-studied Saccharomyces cerevisiae genes with multiple literature references, and the second contains newly discovered genes in Drosophila melanogaster; many have no literature references at all. In both cases, we are able to rapidly define and identify the biologically relevant gene expression profiles without manual intervention. In both cases, we identified novel clusters that were not noted by the original investigators.

Animals↗

A literature-based method for assessing the functional coherence of a gene group.

MOTIVATION: Many experimental and algorithmic approaches in biology generate groups of genes that need to be examined for related functional properties. For example, gene expression profiles are frequently organized into clusters of genes that may share functional properties. We evaluate a method, neighbor divergence per gene (NDPG), that uses scientific literature to assess whether a group of genes are functionally related. The method requires only a corpus of documents and an index connecting the documents to genes. RESULTS: We evaluate NDPG on 2796 functional groups generated by the Gene Ontology consortium in four organisms: mouse, fly, worm and yeast. NDPG finds functional coherence in 96, 92, 82 and 45% of the groups (at 99.9% specificity) in yeast, mouse, fly and worm respectively.

Abstracting and Indexing↗

Identification of osteopontin as a prognostic plasma marker for head and neck squamous cell carcinomas.

PURPOSE: Tumor hypoxia modifies treatment efficacy and promotes tumor progression. Here, we investigated the relationship between osteopontin (OPN), tumor pO(2), and prognosis in patients with head and neck squamous cell carcinomas (HNSCC). EXPERIMENTAL DESIGN: We performed linear discriminant analysis, a machine learning algorithm, on the NCI-60 cancer cell line microarray expression database to identify a gene profile that best distinguish cell lines with high Von-Hippel Lindau (VHL) gene expression, an important regulator of hypoxia-related genes, from those with low expression. Plasma OPN levels in 15 volunteers, 31 VHL patients, and 54 HNSCC patients were quantitatively measured by ELISA. The relationships between plasma OPN levels, tumor pO(2) as measured by the Eppendorf microelectrode, freedom from relapse (FFR), and survival in HNSCC patients were evaluated. RESULTS: Microarray analysis indicated that OPN gene expression inversely correlated with that of VHL. These findings were confirmed by Northern blot analysis. ELISA studies and Western blot in a HNSCC cell line demonstrated that hypoxia exposure resulted in increased OPN secretion. Patients with VHL syndrome had significantly higher plasma OPN levels than healthy volunteers. Plasma OPN level inversely correlated with tumor pO(2) (P = 0.003, r = -0.42). OPN levels correlated with clinical outcomes. The 1-year FFR and survival rates were 80 and 100%, respectively, for patients with OPN levels 450 ng/ml (P = 0.002 and 0.0005). Multivariate analysis revealed that OPN was an independent predictor for FFR and survival. CONCLUSIONS: Plasma OPN levels appeared to correlate with tumor hypoxia in HNSCC patients and may serve as noninvasive tests to identify patients at high risk for tumor recurrence.

Adult↗

Using text analysis to identify functionally coherent gene groups.

The analysis of large-scale genomic information (such as sequence data or expression patterns) frequently involves grouping genes on the basis of common experimental features. Often, as with gene expression clustering, there are too many groups to easily identify the functionally relevant ones. One valuable source of information about gene function is the published literature. We present a method, neighbor divergence, for assessing whether the genes within a group share a common biological function based on their associated scientific literature. The method uses statistical natural language processing techniques to interpret biological text. It requires only a corpus of documents relevant to the genes being studied (e.g., all genes in an organism) and an index connecting the documents to appropriate genes. Given a group of genes, neighbor divergence assigns a numerical score indicating how "functionally coherent" the gene group is from the perspective of the published literature. We evaluate our method by testing its ability to distinguish 19 known functional gene groups from 1900 randomly assembled groups. Neighbor divergence achieves 79% sensitivity at 100% specificity, comparing favorably to other tested methods. We also apply neighbor divergence to previously published gene expression clusters to assess its ability to recognize gene groups that had been manually identified as representative of a common function.

Algorithms↗

Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.

Functional characterizations of thousands of gene products from many species are described in the published literature. These discussions are extremely valuable for characterizing the functions not only of these gene products, but also of their homologs in other organisms. The Gene Ontology (GO) is an effort to create a controlled terminology for labeling gene functions in a more precise, reliable, computer-readable manner. Currently, the best annotations of gene function with the GO are performed by highly trained biologists who read the literature and select appropriate codes. In this study, we explored the possibility that statistical natural language processing techniques can be used to assign GO codes. We compared three document classification methods (maximum entropy modeling, naïve Bayes classification, and nearest-neighbor classification) to the problem of associating a set of GO codes (for biological process) to literature abstracts and thus to the genes associated with the abstracts. We showed that maximum entropy modeling outperforms the other methods and achieves an accuracy of 72% when ascertaining the function discussed within an abstract. The maximum entropy method provides confidence measures that correlate well with performance. We conclude that statistical methods may be used to assign GO codes and may be useful for the difficult task of reassignment as terminology standards evolve over time.

Algorithms↗

Determining the genomic locations of repetitive DNA sequences with a whole-genome microarray: IS6110 in Mycobacterium tuberculosis.

The mycobacterial insertion sequence IS6110 has been exploited extensively as a clonal marker in molecular epidemiologic studies of tuberculosis. In addition, it has been hypothesized that this element is an important driving force behind genotypic variability that may have phenotypic consequences. We present here a novel, DNA microarray-based methodology, designated SiteMapping, that simultaneously maps the locations and orientations of multiple copies of IS6110 within the genome. To investigate the sensitivity, accuracy, and limitations of the technique, it was applied to eight Mycobacterium tuberculosis strains for which complete or partial IS6110 insertion site information had been determined previously. SiteMapping correctly located 64% (38 of 59) of the IS6110 copies predicted by restriction fragment length polymorphism analysis. The technique is highly specific; 97% of the predicted insertion sites were true insertions. Eight previously unknown insertions were identified and confirmed by PCR or sequencing. The performance could be improved by modifications in the experimental protocol and in the approach to data analysis. SiteMapping has general applicability and demonstrates an expansion in the applications of microarrays that complements conventional approaches in the study of genome architecture.

DNA Transposable Elements↗