Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “unsupervised clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Robust dysregulation of gene expression in substantia nigra and striatum in Parkinson's disease.

Large-scale genomics approaches are now widely utilized to study a myriad of human diseases. These powerful techniques, when combined with data analysis tools, detect changes in transcript abundance in diseased tissue relative to control. We hypothesize that specific differential gene expression underlies important pathogenic processes in Parkinson's disease, which is characterized by the gradual loss of dopaminergic neurons in the substantia nigra and consequent loss of dopamine in the striatum. We have therefore examined gene expression levels in the human parkinsonian nigrostriatal pathway, and compared them with those of neurologically normal controls. Using unsupervised clustering methods, we demonstrate that relatively few genes' expression levels can effectively distinguish between disease and control brains. Further, we identify several interesting patterns of gene expression that illuminate pathogenic cascades in Parkinson's disease. In particular is the robust loss of synaptic gene expression in diseased substantia nigra and striatum.

Aged↗

Intelligent initialization of resource allocating RBF networks.

In any neural network system, proper parameter initialization reduces training time and effort, and generally leads to compact modeling of the process under examination, i.e. less complex network structures and better generalization. However, in cases of multi-dimensional data, parameter initialization is both difficult and time consuming. In the proposed scheme a novel, multi-dimensional, unsupervised clustering method is used to properly initialize neural network architectures, focusing on resource allocating networks (RAN); both the hidden and output layer parameters are determined by the output of the clustering process, without the need for any user interference. The main contribution of this work is that the proposed approach leads to network structures that are compact, efficient and achieve best classification results, without the need for manual selection of suitable initial network parameters. The efficiency of the proposed method has been tested on several classes of publicly available data, such as iris, Wisconsin and ionosphere data.

Algorithms↗

Integrated single-cell and bulk transcriptomic analysis identifies a novel senescent fibroblast subtype associated with poor prognosis in acral melanoma.

BACKGROUND: Acral melanoma (AM) exhibits significant intratumoral heterogeneity, but its tumor microenvironment (TME) and immune regulation remain unclear. This study aims to dissect TME heterogeneity and establish a prognostic model based on key cell subpopulations. METHODS: We collected AM single-cell RNA sequencing (scRNA-seq) and bulk RNA-seq data from the Gene Expression Omnibus (GEO) and the Cancer Genome Atlas (TCGA). Unsupervised clustering, CellChat, and Scissor analysis were performed to characterize cellular heterogeneity, cell-cell communication, and prognosis-related cell subpopulations. Kaplan-Meier analysis was used to assess the prognostic value of key genes, which were further validated by multiplex immunohistochemistry (mIHC). RESULTS: In AM, Mel_C2, C7, and C9 with high SEMA6A and KIT expression were strongly linked to poor prognosis. We further identified a senescent fibroblast subpopulation (sCAF_CDKN2A) characterized by high fibroblast senescence signature (FSS) scores. Integrating Scissor analysis of fibroblast subtypes with bulk prognostic data, we identified COL3A1, VCAN, and KIT as prognosis-associated genes upregulated in poor-outcome-related fibroblast subsets. Cell-cell communication analysis revealed that sCAF_CDKN2A engages in an immunosuppressive network, interacting with regulatory T cells (Tregs) via MIF signaling and receiving signals from exhausted CD8+ T cells through PPIA-BSG interactions. Using transcription factor expression patterns from these fibroblast subtypes, we constructed a prognostic model that effectively stratified patients into distinct risk groups with significant differences in overall survival (OS). mIHC confirmed significantly higher protein levels of SEMA6A and COL3A1 in tumor tissues compared to matched normal tissues. CONCLUSIONS: We established a novel prognostic model for AM and identified sCAF_CDKN2A as an immunosuppressive senescent fibroblast subpopulation driving poor prognosis.

Acral melanoma↗

Interpreting expression profiles of cancers by genome-wide survey of breadth of expression in normal tissues.

A critical and difficult part of studying cancer with DNA microarrays is data interpretation. Besides the need for data analysis algorithms, integration of additional information about genes might be useful. We performed genome-wide expression profiling of 36 types of normal human tissues and identified 2503 tissue-specific genes. We then systematically studied the expression of these genes in cancers by reanalyzing a large collection of published DNA microarray datasets. We observed that the expression level of liver-specific genes in hepatocellular carcinoma (HCC) correlates with the clinically defined degree of tumor differentiation. Through unsupervised clustering of tissue-specific genes differentially expressed in tumors, we extracted expression patterns that are characteristic of individual cell types, uncovering differences in cell lineage among tumor subtypes. We were able to detect the expression signature of hepatocytes in HCC, neuron cells in medulloblastoma, glia cells in glioma, basal and luminal epithelial cells in breast tumors, and various cell types in lung cancer samples. We also demonstrated that tissue-specific expression signatures are useful in locating the origin of metastatic tumors. Our study shows that integration of each gene's breadth of expression (BOE) in normal tissues is important for biological interpretation of the expression profiles of cancers in terms of tumor differentiation, cell lineage, and metastasis.

Algorithms↗

Identification of novel candidate oncogenes and tumor suppressors in malignant pleural mesothelioma using large-scale transcriptional profiling.

Malignant pleural mesothelioma (MPM) is a highly lethal, poorly understood neoplasm that is typically associated with asbestos exposure. We performed transcriptional profiling using high-density oligonucleotide microarrays containing approximately 22,000 genes to elucidate potential molecular and pathobiological pathways in MPM using discarded human MPM tumor specimens (n = 40), normal lung specimens (n = 4), normal pleura specimens (n = 5), and MPM and SV40-immortalized mesothelial cell lines (n = 5). In global expression analysis using unsupervised clustering techniques, we found two potential subclasses of mesothelioma that correlated loosely with tumor histology. We also identified sets of genes with expression levels that distinguish between multiple tumor subclasses, normal and tumor tissues, and tumors with different morphologies. Microarray gene expression data were confirmed using quantitative reverse transcriptase-polymerase chain reaction and protein analysis for three novel candidate oncogenes (NME2, CRI1, and PDGFC) and one candidate tumor suppressor (GSN). Finally, we used bioinformatics tools (ie, software) to create and explore complex physiological pathways. Combined, all of these data may advance our understanding of mesothelioma tumorigenesis, pathobiology, or both.

Blotting, Western↗

Metabolic fingerprinting of salt-stressed tomatoes.

The aim of this study was to adopt the approach of metabolic fingerprinting through the use of Fourier transform infrared (FT-IR) spectroscopy and chemometrics to study the effect of salinity on tomato fruit. Two varieties of tomato were studied, Edkawy and Simge F1. Salinity treatment significantly reduced the relative growth rate of Simge F1 but had no significant effect on that of Edkawy. In both tomato varieties salt-treatment significantly reduced mean fruit fresh weight and size class but had no significant affect on total fruit number. Marketable yield was however reduced in both varieties due to the occurrence of blossom end rot in response to salinity. Whole fruit flesh extracts from control and salt-grown tomatoes were analysed using FT-IR spectroscopy. Each sample spectrum contained 882 variables, absorbance values at different wavenumbers, making visual analysis difficult and therefore machine learning methods were applied. The unsupervised clustering method, principal component analysis (PCA) showed no discrimination between the control and salt-treated fruit for either variety. The supervised method, discriminant function analysis (DFA) was able to classify control and salt-treated fruit in both varieties. Genetic algorithms (GA) were applied to identify discriminatory regions within the FT-IR spectra important for fruit classification. The GA models were able to classify control and salt-treated fruit with a typical error, when classifying the whole data set, of 9% in Edkawy and 5% in Simge F1. Key regions were identified within the spectra corresponding to nitrile containing compounds and amino radicals. The application of GA enabled the identification of functional groups of potential importance in relation to the response of tomato to salinity.

Algorithms↗

Automated interpretation of protein subcellular location patterns.

Proteomics is a major current focus of biomedical research, and location proteomics is the important branch of proteomics that systematically studies the subcellular distributions for all proteins expressed in a given cell type. Fluorescence microscopy of labeled proteins is currently the main methodology to obtain location information. Traditionally, microscope images are analyzed by visual inspection, which suffers from inefficiency and inconsistency. Automated and objective interpretation approaches are therefore needed for location proteomics. In this article, we briefly review recent advances in automated imaging interpretation tools, including supervised classification (which assigns location pattern labels to previously unseen images), unsupervised clustering (which groups proteins based on the similarity among their subcellular distributions), and additional statistical tools that can aid cell and molecular biologists who use microscopy in their work.

Automation↗

Methylation profiling of normal tissue adjacent to breast tumors reveals two distinct groups with divergent tumor microenvironment features.

We previously identified diverse genetic evolutionary patterns in whole-genome sequencing of paired normal tissue adjacent to tumor (NAT) and tumor tissues from Hong Kong breast cancer (HKBC) patients. Here, we investigated whether DNA methylation (DNAm) contributes to NAT heterogeneity and shapes the tumor microenvironment (TME). Genome-wide DNAm profiling was performed on paired NAT and tumor tissues from 188 HKBC patients using the Infinium 850 K array. RNA-seq data were available for 76 NATs and 177 tumors. Cellular composition was inferred using MethylCIBERSORT, CIBERSORTx, and EpiDISH, and histopathologic features were assessed on 115 H&E-stained sections. Unsupervised clustering identified two distinct NAT subtypes with divergent TME characteristics. Cluster 1 (N = 139) showed higher epithelial and fibroblast content and enrichment of estrogen response pathways. Cluster 2 (N = 49) exhibited an immune-metabolic phenotype characterized by increased fat and immune cells, stromal disruption, inflammatory pathway activation, and greater macrophage infiltration. Cluster 2 patients also demonstrated significantly younger epigenetic age estimated using multiple epigenetic clocks. These DNAm-defined NAT subtypes and associated TME features were validated in 97 NAT samples from TCGA breast cancer patients. Overall, our findings identify DNAm-driven NAT heterogeneity with distinct TME landscapes, providing new insights into field cancerization and tumor evolution in breast cancer.

Journal Article↗

Quantitative essentiality in a reduced genome: a functional, regulatory and structural fitness map.

Essentiality studies have traditionally focused on coding regions, often overlooking other small genetic regulatory elements. To address this, we combined transposon libraries containing promoter or terminator sequences to obtain a high-resolution essentiality map of a genome-reduced bacterium, at near-single-nucleotide precision when considering non-essential genes. By integrating temporal transposon-sequencing data by k-means unsupervised clustering, we present a novel essentiality assessment approach, providing dynamic and quantitative information on the fitness contribution of different genomic regions. We compared the insertion tolerance and persistence of the two engineered libraries, assessing the local impact of transcription and termination on cell fitness. Essentiality assessment at the local base-level revealed essential protein domains and small genomic regions that are either essential or inaccessible to transposon insertion. We also identified structural regions within essential genes that tolerate transposon disruptions, resulting in functionally split proteins. Overall, this study presents a nuanced view of gene essentiality, shifting from static and binary models to a more accurate perspective. Additionally, it provides valuable insights for genome engineering and enhances our understanding of the biology of genome-reduced cells.

DNA Transposable Elements↗

Gene expression profiling of primary cultures of ovarian epithelial cells identifies novel molecular classifiers of ovarian cancer.

In order to elucidate the biological variance between normal ovarian surface epithelial (NOSE) and epithelial ovarian cancer (EOC) cells, and to build a molecular classifier to discover new markers distinguishing these cells, we analysed gene expression patterns of 65 primary cultures of these tissues by oligonucleotide microarray. Unsupervised clustering highlights three subgroups of tumours: low malignant potential tumours, invasive solid tumours and tumour cells derived from ascites. We selected 18 genes with expression profiles that enable the distinction of NOSE from these three groups of EOC with 92% accuracy. Validation using an independent published data set derived from tissues or primary cultures confirmed a high accuracy (87-96%). The distinctive expression pattern of a subset of genes was validated by quantitative reverse transcription-PCR. An ovarian-specific tissue array representing tissues from NOSE and EOC samples of various subtypes and grades was used to further assess the protein expression patterns of two differentially expressed genes (Msln and BMP-2) by immunohistochemistry. This study highlights the relevance of using primary cultures of epithelial ovarian cells as a model system for gene profiling studies and demonstrates that the statistical analysis of gene expression profiling is a useful approach for selecting novel molecular tumour markers.

Biomarkers, Tumor↗

Gene expression profiling of human ovarian tumours.

There is currently a lack of reliable diagnostic and prognostic markers for ovarian cancer. We established gene expression profiles for 120 human ovarian tumours to identify determinants of histologic subtype, grade and degree of malignancy. Unsupervised cluster analysis of the most variable set of expression data resulted in three major tumour groups. One consisted predominantly of benign tumours, one contained mostly malignant tumours, and one was comprised of a mixture of borderline and malignant tumours. Using two supervised approaches, we identified a set of genes that distinguished the benign, borderline and malignant phenotypes. These algorithms were unable to establish profiles for histologic subtype or grade. To validate these findings, the expression of 21 candidate genes selected from these analyses was measured by quantitative RT-PCR using an independent set of tumour samples. Hierarchical clustering of these data resulted in two major groups, one benign and one malignant, with the borderline tumours interspersed between the two groups. These results indicate that borderline ovarian tumours may be classified as either benign or malignant, and that this classifier could be useful for predicting the clinical course of borderline tumours. Immunohistochemical analysis also demonstrated increased expression of CD24 antigen in malignant versus benign tumour tissue. The data that we have generated will contribute to a growing body of expression data that more accurately define the biologic and clinical characteristics of ovarian cancers.

Adenocarcinoma, Clear Cell↗

Comparison of gene expression profiling between malignant and normal plasma cells with oligonucleotide arrays.

The DNA microarray technology enables the identification of the large number of genes involved in the complex deregulation of cell homeostasis taking place in cancer. Using Affymetrix microarrays, we have compared the gene expression profiles of highly purified malignant plasma cells from nine patients with multiple myeloma (MM) and eight myeloma cell lines to those of highly purified nonmalignant plasma cells (eight samples) obtained by in vitro differentiation of peripheral blood B cells. Two unsupervised clustering algorithms classified these 25 samples into two distinct clusters: a malignant plasma cell cluster and a normal plasma cell cluster. Two hundred and fifty genes were significantly up-regulated and 159 down-regulated in malignant plasma samples compared to normal plasma samples. For some of these genes, an overexpression or downregulation of the encoded protein was confirmed (cyclin D1, c-myc, BMI-1, cystatin c, SPARC, RB). Two genes overexpressed in myeloma cells (ABL and cystathionine beta synthase) code for enzymes that could be a therapeutic target with specific drugs. These data provide a new insight into the understanding of myeloma disease and prefigure that the development of DNA microarray could help to develop an 'à la carte' treatment in cancer disease.

Adult↗

Identification of radiation-specific responses from gene expression profile.

The responses to ionizing radiation (IR) in tumors are dependent on cellular context. We investigated radiation-related expression patterns in Jurkat T cells with nonsense mutation in p53 using cDNA microarray. Expression of 2400 genes in gamma-irradiated cells was distinct from other stimulations like anti-CD3, phetohemagglutinin (PHA) and concanavalin A (ConA) in unsupervised clustering analysis. Among them, 384 genes were selected for their IR-specific changes to make 'RadChip'. In spite of p53 status, every type of cells showed similar patterns in expression of these genes upon gamma-radiation. Moreover, radiation-induced responses were clearly separated from the responses to other genotoxic stress like UV radiation, cisplatin and doxorubicin. We focused on two IR-related genes, phospholipase Cgamma2 (PLCG2) and cytosolic epoxide hydrolase (EPHX2), which were increased at 12 h after gamma-radiation in RT-PCR. TPCK could suppress the induction of these two genes in either of Jurkat T cells and PBMCs, which might suggest the transcriptional regulation of PLCG2 and EPHX2 by NF-kappaB upon gamma-radiation. From these results, we could identify the IR-specific genes from expression profiling, which can be used as radiation biomarkers to screen radiation exposure as well as probing the mechanism of cellular responses to ionizing radiation.

Apoptosis↗

Expression profiling reveals a distinct transcription signature in follicular thyroid carcinomas with a PAX8-PPAR(gamma) fusion oncogene.

The demonstration of the PAX8-PPAR(gamma) fusion oncogene in a subset of follicular thyroid tumors provides a new and promising starting point to dissect the molecular genetic events involved in the development of this tumor form. In the present study, we compared the gene expression profiles of follicular thyroid carcinomas (FTCs) bearing a PAX8-PPAR(gamma) fusion against FTCs that lack this fusion. Using unsupervised clustering and multidimensional scaling analyses, we show that FTCs possessing a PAX8-PPAR(gamma) fusion have a highly uniform and distinct gene expression signature that clearly distinguishes them from FTCs without the fusion. The PAX8-PPAR(gamma)(+) FTCs grouped in a defined cluster, where highly ranked genes were mostly associated with signal transduction, cell growth and translation control. Notably, a large number of ribosomal protein and translation-associated genes were concurrently underexpressed in the FTCs with the fusion. Taken together, our findings further support that follicular carcinomas with a PAX8-PPAR(gamma) rearrangement constitute a distinct biological entity. The current data represent one step to elucidate the molecular pathways in the development of FTCs with the specific PAX8-PPAR(gamma) fusion.

Adenocarcinoma, Follicular↗

Supervised classification for gene network reconstruction.

One of the central problems of functional genomics is revealing gene expression networks - the relationships between genes that reflect observations of how the expression level of each gene affects those of others. Microarray data are currently a major source of information about the interplay of biochemical network participants in living cells. Various mathematical techniques, such as differential equations, Bayesian and Boolean models and several statistical methods, have been applied to expression data in attempts to extract the underlying knowledge. Unsupervised clustering methods are often considered as the necessary first step in visualization and analysis of the expression data. As for supervised classification, the problem mainly addressed so far has been how to find discriminative genes separating various samples or experimental conditions. Numerous methods have been applied to identify genes that help to predict treatment outcome or to confirm a diagnosis, as well as to identify primary elements of gene regulatory circuits. However, less attention has been devoted to using supervised learning to uncover relationships between genes and/or their products. To start filling this gap a machine-learning approach for gene networks reconstruction is described here. This approach is based on building classifiers--functions, which determine the state of a gene's transcription machinery through expression levels of other genes. The method can be applied to various cases where relationships between gene expression levels could be expected.

Genes↗

Gene expression profiles of breast cancer obtained from core cut biopsies before neoadjuvant docetaxel, adriamycin, and cyclophoshamide chemotherapy correlate with routine prognostic markers and could be used to identify predictive signatures.

BACKGROUND: Neoadjuvant administration of chemotherapy provides a unique opportunity to monitor response to treatment in breast cancer and assesses response exactly. Global gene expression profiling by microarrays has been used as a valuable tool for the identification of prognostic and predictive marker genes. Even though this technology is now wide spread and relatively standardized, there are only few data available which compare established parameters with expression values to determine reliability of this method. Therefore we analyzed gene expression data of pretreatment biopsies of breast cancer patients and compared them with the results of the immunohistochemical receptor expression for ER/ PR and Her-2, as well as FISH testing for HER-2 amplification. We analyzed the change of expression of these markers before and after neoadjuvant chemotherapy. Furthermore we evaluated the predictive significance of prognostic gene signatures as described by Sorlie, van't Veer and Ahr for response to neoadjuvant chemotherapy. METHODS: Pretherapeutic core biopsies were obtained from 70 patients undergoing neoadjuvant TAC chemotherapy within the GEPARTRIO-trial. Samples were characterized according to standard pathology including ER, PR and HER2 IHC and amount of cancer cells. Only biopsies with more than 80 % tumor cells were considered for further examination. RNA was isolated and expression profiling performed using Affymetrix Hg U133 Arrays (22 500 genes). GeneData's Expressionist software was used for bioinformatic analyses. RESULTS: More than two thirds of the biopsies yielded sufficient amounts (> 5 microg) of RNA for expression profiling and high quality data were obtained for 50 samples. Unsupervised clustering broadly revealed a correlation with hormone receptor status. When ER-alpha, PR and HER2 as analyzed by immunohistochemistry were compared to the corresponding mRNA data from gene chips more than 90 % concordance was observed. We could observe a switch of receptor expression for ER, PR or HER-2 from positive to negative and vice versa in 16/35 cases (45.7 %) and 5/22 cases (22.7 %) respectively. The prognostic marker sets of Sorlie, van't Veer and Ahr could not discriminate responders from non-responders in our patient group. CONCLUSIONS: Our results demonstrate that reliable expression profiles can be achieved by using limited amounts of tissue obtained during neoadjuvant chemotherapy. Microarray data capture conventional prognostic markers but might contain additional informative gene sets correlated with treatment outcome. Prognostic marker sets are not suitable to predict tumor response in the neoadjuvant setting, suggesting the necessity of class prediction methods to identify marker sets predictive for the type of therapy used.

Adult↗

Prognostically useful gene-expression profiles in acute myeloid leukemia.

BACKGROUND: In patients with acute myeloid leukemia (AML) a combination of methods must be used to classify the disease, make therapeutic decisions, and determine the prognosis. However, this combined approach provides correct therapeutic and prognostic information in only 50 percent of cases. METHODS: We determined the gene-expression profiles in samples of peripheral blood or bone marrow from 285 patients with AML using Affymetrix U133A GeneChips containing approximately 13,000 unique genes or expression-signature tags. Data analyses were carried out with Omniviz, significance analysis of microarrays, and prediction analysis of microarrays software. Statistical analyses were performed to determine the prognostic significance of cases of AML with specific molecular signatures. RESULTS: Unsupervised cluster analyses identified 16 groups of patients with AML on the basis of molecular signatures. We identified the genes that defined these clusters and determined the minimal numbers of genes needed to identify prognostically important clusters with a high degree of accuracy. The clustering was driven by the presence of chromosomal lesions (e.g., t(8;21), t(15;17), and inv(16)), particular genetic mutations (CEBPA), and abnormal oncogene expression (EVI1). We identified several novel clusters, some consisting of specimens with normal karyotypes. A unique cluster with a distinctive gene-expression signature included cases of AML with a poor treatment outcome. CONCLUSIONS: Gene-expression profiling allows a comprehensive classification of AML that includes previously identified genetically defined subgroups and a novel cluster with an adverse prognosis.

Acute Disease↗

Knowledge-based analysis of microarray gene expression data by using support vector machines.

We introduce a method of functionally classifying genes by using gene expression data from DNA microarray hybridization experiments. The method is based on the theory of support vector machines (SVMs). SVMs are considered a supervised computer learning method because they exploit prior knowledge of gene function to identify unknown genes of similar function from expression data. SVMs avoid several problems associated with unsupervised clustering methods, such as hierarchical clustering and self-organizing maps. SVMs have many mathematical features that make them attractive for gene expression analysis, including their flexibility in choosing a similarity function, sparseness of solution when dealing with large data sets, the ability to handle large feature spaces, and the ability to identify outliers. We test several SVMs that use different similarity metrics, as well as some other supervised learning methods, and find that the SVMs best identify sets of genes with a common function using expression data. Finally, we use SVMs to predict functional roles for uncharacterized yeast ORFs based on their expression data.

Algorithms↗