Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “unsupervised clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Microarray data analysis: from hypotheses to conclusions using gene expression data.

We review several commonly used methods for the design and analysis of microarray data. To begin with, some experimental design issues are addressed. Several approaches for pre-processing the data (filtering and normalization) before the statistical analysis stage are then discussed. A common first step in this type of analysis is gene selection based on statistical testing. Two approaches, permutation and model-based methods are explained and we emphasize the need to correct for multiple testing. Moreover, powerful approaches based on gene sets are mentioned. Clustering of either genes or samples is frequently performed when analyzing microarray data. We summarize the basics of both supervised and unsupervised clustering (classification). The latter may be of use for creating diagnostic arrays, for example. Construction of biological networks, such as pathways, is a statistically challenging but complex task that is a relatively new development and hence mentioned only briefly. We finish with some remarks on literature and software. The emphasis in this paper is on the philosophy behind several statistical issues and on a critical interpretation of microarray related analysis methods.

Cluster Analysis↗

Classification of projection neurons and interneurons in the rat lateral amygdala based upon cluster analysis.

Neurons in the rat lateral amygdala in situ were classified based upon electrophysiological and molecular parameters, as studied by patch-clamp, single-cell RT-PCR and unsupervised cluster analyses. Projection neurons (class I) were characterized by low firing rates, frequency adaptation and expression of the vesicular glutamate transporter (VGLUT1). Two classes were distinguished based upon electrotonic properties and the presence (IB) or absence (IA) of vasointestinal peptide (VIP). Four classes of glutamate decarboxylase (GAD67) containing interneurons were encountered. Class III reflected "classical" interneurons, generating fast spikes with no frequency adaptation. Class II neurons generated fast spikes with early frequency adaptation and differed from class III by the presence of VIP and the relatively rare presence of neuropeptide Y (NPY) and somatostatin (SOM). Class IV and V were not clearly separated by molecular markers, but by membrane potential values and spike patterns. Morphologically, projection neurons were large, spiny cells, whereas the other neuronal classes displayed smaller somata and spine-sparse dendrites.

Amygdala↗

Cell-nuclear data reduction and prognostic model selection in bladder tumor recurrence.

OBJECTIVE: The paper aims at improving the prediction of superficial bladder recurrence. To this end, feedforward neural networks (FNNs) and a feature selection method based on unsupervised clustering, were employed. MATERIAL AND METHODS: A retrospective prognostic study of 127 patients diagnosed with superficial urinary bladder cancer was performed. Images from biopsies were digitized and cell nuclei features were extracted. To design FNN classifiers, different training methods and architectures were investigated. The unsupervised k-windows (UKW) and the fuzzy c-means clustering algorithms were applied on the feature set to identify the most informative feature subsets. RESULTS: UKW managed to reduce the dimensionality of the feature space significantly, and yielded prediction rates 87.95% and 91.41%, for non-recurrent and recurrent cases, respectively. The prediction rates achieved with the reduced feature set were marginally lower compared to the ones attained with the complete feature set. The training algorithm that exhibited the best performance in all cases was the adaptive on-line backpropagation algorithm. CONCLUSIONS: FNNs can contribute to the accurate prognosis of bladder cancer recurrence. The proposed feature selection method can remove redundant information without a significant loss in predictive accuracy, and thereby render the prognostic model less complex, more robust, and hence suitable for clinical use.

Algorithms↗

Visualization of multiple influences on ocellar flight control in giant honeybees with the data-mining tool Viscovery SOMine.

Viscovery SOMine is a software tool for advanced analysis and monitoring of numerical data sets. It was developed for professional use in business, industry, and science and to support dependency analysis, deviation detection, unsupervised clustering, nonlinear regression, data association, pattern recognition, and animated monitoring. Based on the concept of self-organizing maps (SOMs), it employs a robust variant of unsupervised neural networks--namely, Kohonen's Batch-SOM, which is further enhanced with a new scaling technique for speeding up the learning process. This tool provides a powerful means by which to analyze complex data sets without prior statistical knowledge. The data representation contained in the trained SOM is systematically converted to be used in a spectrum of visualization techniques, such as evaluating dependencies between components, investigating geometric properties of the data distribution, searching for clusters, or monitoring new data. We have used this software tool to analyze and visualize multiple influences of the ocellar system on free-flight behavior in giant honeybees. Occlusion of ocelli will affect orienting reactivities in relation to flight target, level of disturbance, and position of the bee in the flight chamber; it will induce phototaxis and make orienting imprecise and dependent on motivational settings. Ocelli permit the adjustment of orienting strategies to environmental demands by enforcing abilities such as centering or flight kinetics and by providing independent control of posture and flight course.

Animals↗

Unsupervised immunophenotypic profiling of chronic lymphocytic leukemia.

BACKGROUND: Proteomics and functional genomics have revolutionized approaches to disease classification. Like proteomics, flow cytometry (FCM) assesses concurrent expression of many proteins, with the advantage of using intact cells that may be differentially selected during analysis. However, FCM has generally been used for incremental marker validation or construction of predictive models based on known patterns, rather than as a tool for unsupervised class discovery. We undertook a retrospective analysis of clinical FCM data to assess the feasibility of a cell-based proteomic approach to FCM by unsupervised cluster analysis. METHODS: Multicolor FCM data on peripheral blood (PB) and bone marrow (BM) lymphocytes from 140 consecutive patients with B-cell chronic lymphoproliferative disorders (LPDs), including 81 chronic lymphocytic leukemia (CLLs), were studied. Expression was normalized for CD19 totals, and recorded for 10 additional B-cell markers. Data were subjected to hierarchical cluster analysis using complete linkage by Pearson's correlation. Analysis of CLL in PB samples (n = 63) discovered three major clusters. One cluster (14 patients) was skewed toward "atypical" CLL and was characterized by high CD20, CD22, FMC7, and light chain, and low CD23. The remaining two clusters consisted almost entirely (48/49) of cases recorded as typical BCLL. The smaller "typical" BCLL cluster differed from the larger cluster by high CD38 (P = 0.001), low CD20 (P = 0.001), and low CD23 (P = 0.016). These two typical BCLL clusters showed a trend toward a difference in survival (P = 0.1090). Statistically significant cluster stability was demonstrated by expanding the dataset to include BM samples, and by using a method of random sampling with replacement. CONCLUSIONS: This study supports the concept that unsupervised immunophenotypic profiling of FCM data can yield reproducible subtypes of lymphoma/chronic leukemia. Expanded studies are warranted in the use of FCM as an unsupervised class discovery tool, akin to other proteomic methods, rather than as a validation tool.

ADP-ribosyl Cyclase 1↗

Supervised cluster analysis for microarray data based on multivariate Gaussian mixture.

MOTIVATION: Grouping genes having similar expression patterns is called gene clustering, which has been proved to be a useful tool for extracting underlying biological information of gene expression data. Many clustering procedures have shown success in microarray gene clustering; most of them belong to the family of heuristic clustering algorithms. Model-based algorithms are alternative clustering algorithms, which are based on the assumption that the whole set of microarray data is a finite mixture of a certain type of distributions with different parameters. Application of the model-based algorithms to unsupervised clustering has been reported. Here, for the first time, we demonstrated the use of the model-based algorithm in supervised clustering of microarray data. RESULTS: We applied the proposed methods to real gene expression data and simulated data. We showed that the supervised model-based algorithm is superior over the unsupervised method and the support vector machines (SVM) method. AVAILABILITY: The program written in the SAS language implementing methods I-III in this report is available upon request. The software of SVMs is available in the website http://svm.sdsc.edu/cgi-bin/nph-SVMsubmit.cgi

Algorithms↗

Mass distributed clustering: a new algorithm for repeated measurements in gene expression data.

The availability of whole-genome sequence data and high-throughput techniques such as DNA microarray enable researchers to monitor the alteration of gene expression by a certain organ or tissue in a comprehensive manner. The quantity of gene expression data can be greater than 30,000 genes per one measurement, making data clustering methods for analysis essential. Biologists usually design experimental protocols so that statistical significance can be evaluated; often, they conduct experiments in triplicate to generate a mean and standard deviation. Existing clustering methods usually use these mean or median values, rather than the original data, and take significance into account by omitting data showing large standard deviations, which eliminates potentially useful information. We propose a clustering method that uses each of the triplicate data sets as a probability distribution function instead of pooling data points into a median or mean. This method permits truly unsupervised clustering of the data from DNA microarrays.

Algorithms↗

Gene expression profiling of adult acute myeloid leukemia identifies novel biologic clusters for risk classification and outcome prediction.

To determine whether gene expression profiling could improve risk classification and outcome prediction in older acute myeloid leukemia (AML) patients, expression profiles were obtained in pretreatment leukemic samples from 170 patients whose median age was 65 years. Unsupervised clustering methods were used to classify patients into 6 cluster groups (designated A to F) that varied significantly in rates of resistant disease (RD; P < .001), complete response (CR; P = .023), and disease-free survival (DFS; P = .023). Cluster A (n = 24), dominated by NPM1 mutations (78%), normal karyotypes (75%), and genes associated with signaling and apoptosis, had the best DFS (27%) and overall survival (OS; 25% at 5 years). Patients in clusters B (n = 22) and C (n = 31) had the worst OS (5% and 6%, respectively); cluster B was distinguished by the highest rate of RD (77%) and multidrug resistant gene expression (ABCG2, MDR1). Cluster D was characterized by a "proliferative" gene signature with the highest proportion of detectable cytogenetic abnormalities (76%; including 83% of all favorable and 34% of unfavorable karyotypes). Cluster F (n = 33) was dominated by monocytic leukemias (97% of cases), also showing increased NPM1 mutations (61%). These gene expression signatures provide insights into novel groups of AML not predicted by traditional studies that impact prognosis and potential therapy.

Acute Disease↗

Absence of a specific radiation signature in post-Chernobyl thyroid cancers.

Thyroid cancers have been the main medical consequence of the Chernobyl accident. On the basis of their pathological features and of the fact that a large proportion of them demonstrate RET-PTC translocations, these cancers are considered as similar to classical sporadic papillary carcinomas, although molecular alterations differ between both tumours. We analysed gene expression in post-Chernobyl cancers, sporadic papillary carcinomas and compared to autonomous adenomas used as controls. Unsupervised clustering of these data did not distinguish between the cancers, but separates both cancers from adenomas. No gene signature separating sporadic from post-Chernobyl PTC (chPTC) could be found using supervised and unsupervised classification methods although such a signature is demonstrated for cancers and adenomas. Furthermore, we demonstrate that pooled RNA from sporadic and chPTC are as strongly correlated as two independent sporadic PTC pools, one from Europe, one from the US involving patients not exposed to Chernobyl radiations. This result relies on cDNA and Affymetrix microarrays. Thus, platform-specific artifacts are controlled for. Our findings suggest the absence of a radiation fingerprint in the chPTC and support the concept that post-Chernobyl cancer data, for which the cancer-causing event and its date are known, are a unique source of information to study naturally occurring papillary carcinomas.

Adenoma↗

Spectral clustering algorithms for ultrasound image segmentation.

Image segmentation algorithms derived from spectral clustering analysis rely on the eigenvectors of the Laplacian of a weighted graph obtained from the image. The NCut criterion was previously used for image segmentation in supervised manner. We derive a new strategy for unsupervised image segmentation. This article describes an initial investigation to determine the suitability of such segmentation techniques for ultrasound images. The extension of the NCut technique to the unsupervised clustering is first described. The novel segmentation algorithm is then performed on simulated ultrasound images. Tests are also performed on abdominal and fetal images with the segmentation results compared to manual segmentation. Comparisons with the classical NCut algorithm are also presented. Finally, segmentation results on other types of medical images are shown.

Algorithms↗

Discriminant analysis to evaluate clustering of gene expression data.

In this work we present a procedure that combines classical statistical methods to assess the confidence of gene clusters identified by hierarchical clustering of expression data. This approach was applied to a publicly released Drosophila metamorphosis data set [White et al., Science 286 (1999) 2179-2184]. We have been able to produce reliable classifications of gene groups and genes within the groups by applying unsupervised (cluster analysis), dimension reduction (principal component analysis) and supervised methods (linear discriminant analysis) in a sequential form. This procedure provides a means to select relevant information from microarray data, reducing the number of genes and clusters that require further biological analysis.

Animals↗

Molecular classification of parathyroid neoplasia by gene expression profiling.

The current classification of sporadic parathyroid neoplasia, specifically the distinction of adenoma from multiple gland neoplasia (double adenoma and nonfamilial primary hyperplasia) is problematic and results in a relatively high rate of clinical error. Oligonucleotide microarrays (Affymetrix U133A) were used to evaluate parathyroid samples from 61 patients; 35 adenomas, 10 nonfamilial multiple gland neoplasia, 3 familial primary hyperplasia, 8 renal-induced hyperplasia, and 5 from patients without parathyroid disease (normals). A multiclass comparison using supervised clustering identified distinct gene signatures for each class of parathyroid samples. We developed a predictor model that correctly identified 34 of 35 cases of adenoma, 9 of 10 cases of nonfamilial multiple gland neoplasia, and identified a minimum set of 11 genes for the distinction of adenoma versus multiple gland neoplasia. All methods of unsupervised clustering showed two related but different types of parathyroid adenomas that we have arbitrarily designated as type 1 and type 2 adenomas. Multiple gland parathyroid neoplasia, which represents either synchronous or asynchronous autonomous growth in two, three, or all four parathyroid glands, is a distinct molecular entity and does not represent the molecular pathogenesis of adenoma occurring in multiple glands.

Adenoma↗

Distinguishing key biological pathways between primary breast cancers and their lymph node metastases by gene function-based clustering analysis.

In order to identify key biological pathways that can distinguish between primary breast cancers and their lymph node metastases, we employed gene expression profiling together with gene function-based clustering analysis. We first acquired gene expression profiles of 9 matched primary tumors and the corresponding metastases that contained at least 75% of tumor cells. Then, we applied a clustering algorithm to the preprocessed data. In order to focus on the most informative genes, we ranked all the genes individually based on their abilities to separate the primary breast tumor and metastases samples. Further, we separated these genes into six functional groups according to the Stanford SOURCE database: 'cell cycle,' 'apoptosis,' 'metabolism,' 'cell adhesion and migration,' 'signal transduction,' and 'transcriptional factor and DNA binding molecules.' Unsupervised clustering analysis using all of the 2,303 genes on the microarrays was not able to separate the primary and metastases samples. Clustering analysis using the most informative genes revealed that primary tumors were more tightly clustered, whereas the metastases samples were relatively heterogeneous. The clustering analysis with the genes belonging to different functional groups showed that different functional gene sets varied in their abilities to separate primary tumors and their metastases. Marked separations were found with genes involved in metabolism, signal transduction, cell cycle, and transcriptional factor and DNA binding molecules. In contrast, apoptosis and cell adhesion and migration genes did not provide a clear separation of the two groups of samples. These results suggest that metastatic cells have different metabolism and signal transduction activities, regulated by transcriptional events, from the primary tumor cells. The results also suggest that the altered cell adhesion and migration potentials that are required for tumors to metastasize already exist in the primary tumors as a whole.

Biomarkers, Tumor↗

Automatic segmentation of thalamus from brain MRI integrating fuzzy clustering and dynamic contours.

Thalamus is an important neuro-anatomic structure in the brain. In this paper, an automated method is presented to segment thalamus from magnetic resonance images (MRI). The method is based on a discrete dynamic contour model that consists of vertices and edges connecting adjacent vertices. The model starts from an initial contour and deforms by external and internal forces. Internal forces are calculated from local geometry of the model and external forces are estimated from desired image features such as edges. However, thalamus has low contrast and discontinues edges on MRI, making external force estimation a challenge. The problem is solved using a new algorithm based on fuzzy C-means (FCM) unsupervised clustering, Prewitt edge-finding filter, and morphological operators. In addition, manual definition of the initial contour for the model makes the final segmentation operator-dependent. To eliminate this dependency, new methods are developed for generating the initial contour automatically. The proposed approaches are evaluated and validated by comparing automatic and radiologist's segmentation results and illustrating their agreement.

Algorithms↗

Use of FTIR spectroscopy to distinguish between capsular types and capsular quantities in Streptococcus pneumoniae.

Fourier transform infrared (FTIR) spectroscopy has shown remarkable ability in distinguishing between bacterial species and identifying bacterial colony structures, when used in tandem with methods such as cluster analysis, principal component analysis, or linear discriminant analysis. The present work was aimed to evaluate the potential of FTIR-microscopy (FTIR-MSP) to distinguish between different serotypes and capsular quantities of Streptococcus pneumoniae. In general, the results obtained have consistently proven that the spectral information at the region 900-1,185 cm(-1) was sufficient to distinguish between various pneumococcal serotypes. Moreover, the method was able to differentiate between S. pneumoniae phase variants on the basis of their relative carbohydrate content. The unsupervised cluster analysis of the samples showed differences, not only in the carbohydrate content, but also in the region 1,350-1,480 cm(-1), which is dominated by absorptions due to lipids and phospholipids. This approach proved to be useful for the distinction between S. pneumoniae serotypes and between phase variants, which were shown to acquire different pathogenic capacity.

Bacterial Capsules↗

Disordered semantic representation in schizophrenic temporal cortex revealed by neuromagnetic response patterns.

BACKGROUND: Loosening of associations and thought disruption are key features of schizophrenic psychopathology. Alterations in neural networks underlying this basic abnormality have not yet been sufficiently identified. Previously, we demonstrated that spatio-temporal clustering of magnetic brain responses to pictorial stimuli map categorical representations in temporal cortex. This result has opened the possibility to quantify associative strength within and across semantic categories in schizophrenic patients. We hypothesized that in contrast to controls, schizophrenic patients exhibit disordered representations of semantic categories. METHODS: The spatio-temporal clusters of brain magnetic activities elicited by object pictures related to super-ordinate (flowers, animals, furniture, clothes) and base-level (e.g. tulip, rose, orchid, sunflower) categories were analysed in the source space for the time epochs 170-210 and 210-450 ms following stimulus onset and were compared between 10 schizophrenic patients and 10 control subjects. RESULTS: Spatio-temporal correlations of responses elicited by base-level concepts and the difference of within vs. across super-ordinate categories were distinctly lower in patients than in controls. Additionally, in contrast to the well-defined categorical representation in control subjects, unsupervised clustering indicated poorly defined representation of semantic categories in patients. Within the patient group, distinctiveness of categorical representation in the temporal cortex was positively related to negative symptoms and tended to be inversely related to positive symptoms. CONCLUSION: Schizophrenic patients show a less organized representation of semantic categories in clusters of magnetic brain responses than healthy adults. This atypical neural network architecture may be a correlate of loosening of associations, promoting positive symptoms.

Adult↗

Reference and target region modeling of [11C]-(R)-PK11195 brain studies.

UNLABELLED: PET with [(11)C]-(R)-PK11195 is currently the modality of choice for the in vivo imaging of microglial activation in the human brain. In this work we devised a supervised clustering procedure and a new quantification methodology capable of producing binding potential (BP) estimates quantitatively comparable with those derived from plasma input with robust quantitative implementation at the pixel level. METHODS: The new methodology uses predefined kinetic classes to extract a gray matter reference tissue without specific tracer binding and devoid of spurious signals (in particular, blood pool and muscle). Kinetic classes were derived from an historical database of 12 healthy control subjects and from 3 patients with Huntington's disease. BP estimates were obtained using rank-shaping exponential spectral analysis (RS-ESA) (both plasma and reference input) and the simplified reference tissue model (SRTM). Comparison between plasma- derived BPs and those produced with the new reference methodology was performed using 6 additional healthy control subjects. Reliability of the new methodology was performed on 4 test-retest studies of patients with Alzheimer's disease. RESULTS: The new algorithm selected reference voxels in gray matter tissue avoiding regions with specific binding located, in particular, in the venous and arterial circulation. Using the new reference, BP values obtained using a plasma input and a reference input were in excellent agreement and highly correlated (r = 0.811, P < 10(-5)) when calculated with RS-ESA and less so (r = 0.507, P < 0.005) when SRTM was used. In the production of parametric maps, SRTM was used with the new reference extraction, resulting in test-retest variability (10.6%; mean ICC = 0.878) that was superior to that obtained using the previous unsupervised clustering approach (mean ICC = 0.596). CONCLUSION: Reference region modeling combined with supervised reference tissue extraction produces a robust and reproducible quantitative assessment of [(11)C]-(R)-PK11195 studies in the human brain.

Algorithms↗

MRI tissue characterization of experimental cerebral ischemia in rat.

PURPOSE: To extend the ISODATA image segmentation method to characterize tissue damage in stroke, by generating an MRI score for each tissue that corresponds to its histological damage. MATERIALS AND METHODS: After preprocessing and segmentation (using ISODATA clustering), the proposed method scores tissue regions between 1 and 100. Score 1 is assigned to normal brain matter (white or gray matter), and score 100 to cerebrospinal fluid (CSF). Lesion zones are assigned a score based on their relative levels of similarities to normal brain matter and CSF. To evaluate the method, 15 rats were imaged by a 7T MRI system at one of three time points (acute, subacute, chronic) after MCA occlusion. Then they were killed and their brains were sliced and prepared for histological studies. MRI of two or three slices of each rat brain (using two DWI (b = 400, b = 800), one PDWI, one T2WI, and one T1WI) was performed, and an MRI score between 1 and 100 was determined for each region. Segmented regions were mapped onto the histology images and scored on a scale of 1-10 by an experienced pathologist. The MRI scores were validated by comparison with histology scores. To this end, correlation coefficients between the two scores (MRI and histology) were determined. RESULTS: Experimental results showed excellent correlations between MRI and histology scores at different time points. Depending on the reference tissue (gray matter or white matter) used in the standardization, the correlation coefficients ranged from 0.73 (P < 0.0001) to 0.78 (P < 0.0001) using the entire dataset, including acute, subacute, and chronic time points. This suggests that the proposed multiparametric approach accurately identified and characterized ischemic tissue in a rat model of cerebral ischemia at different stages of stroke evolution. CONCLUSION: The proposed approach scores tissue regions and characterizes them using unsupervised clustering and multiparametric image analysis techniques. The method can be used for a variety of applications in the field of computer-aided diagnosis and treatment, including evaluation of response to treatment. For example, volume changes for different zones of the lesion over time (e.g., tissue recovery) can be evaluated.

Animals↗