Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “unsupervised clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Metabolic discrimination of Catharanthus roseus leaves infected by phytoplasma using 1H-NMR spectroscopy and multivariate data analysis.

A comprehensive metabolomic profiling of Catharanthus roseus L. G. Don infected by 10 types of phytoplasmas was carried out using one-dimensional and two-dimensional NMR spectroscopy followed by principal component analysis (PCA), an unsupervised clustering method requiring no knowledge of the data set and used to reduce the dimensionality of multivariate data while preserving most of the variance within it. With a combination of these techniques, we were able to identify those metabolites that were present in different levels in phytoplasma-infected C. roseus leaves than in healthy ones. The infection by phytoplasma in C. roseus leaves causes an increase of metabolites related to the biosynthetic pathways of phenylpropanoids or terpenoid indole alkaloids: chlorogenic acid, loganic acid, secologanin, and vindoline. Furthermore, higher abundance of Glc, Glu, polyphenols, succinic acid, and Suc were detected in the phytoplasma-infected leaves. The PCA of the (1)H-NMR signals of healthy and phytoplasma-infected C. roseus leaves shows that these metabolites are major discriminating factors to characterize the phytoplasma-infected C. roseus leaves from healthy ones. Based on the NMR and PCA analysis, it might be suggested that the biosynthetic pathway of terpenoid indole alkaloids, together with that of phenylpropanoids, is stimulated by the infection of phytoplasma.

Catharanthus↗

Recognition of temporally changing action potentials in multiunit neural recordings.

We present a method to iteratively train an artificial neural network (ANN) or other supervised pattern classifier in order to adaptively recognize and track temporally changing patterns. This method uses recently acquired data and the existing classifier to create new training sets, from which a new classifier is then trained. The procedure is repeated periodically using the most recently trained classifier. This scheme was evaluated by applying it to simulated situations that arise in chronic recordings of multiunit neural activity from peripheral nerves. The method was able to track the changes in these simulated chronic recordings and to provide better unit recognition rates than an unsupervised clustering method suited to this problem.

Action Potentials↗

Automatic tumor segmentation using knowledge-based techniques.

A system that automatically segments and labels glioblastoma-multiforme tumors in magnetic resonance images (MRI's) of the human brain is presented. The MRI's consist of T1-weighted, proton density, and T2-weighted feature images and are processed by a system which integrates knowledge-based (KB) techniques with multispectral analysis. Initial segmentation is performed by an unsupervised clustering algorithm. The segmented image, along with cluster centers for each class are provided to a rule-based expert system which extracts the intracranial region. Multispectral histogram analysis separates suspected tumor from the rest of the intracranial region, with region analysis used in performing the final tumor labeling. This system has been trained on three volume data sets and tested on thirteen unseen volume data sets acquired from a single MRI system. The KB tumor segmentation was compared with supervised, radiologist-labeled "ground truth" tumor volumes and supervised k-nearest neighbors tumor segmentations. The results of this system generally correspond well to ground truth, both on a per slice basis and more importantly in tracking total tumor volume during treatment over time.

Algorithms↗

Semisupervised learning for molecular profiling.

Class prediction and feature selection are two learning tasks that are strictly paired in the search of molecular profiles from microarray data. Researchers have become aware how easy it is to incur a selection bias effect, and complex validation setups are required to avoid overly optimistic estimates of the predictive accuracy of the models and incorrect gene selections. This paper describes a semisupervised pattern discovery approach that uses the by-products of complete validation studies on experimental setups for gene profiling. In particular, we introduce the study of the patterns of single sample responses (sample-tracking profiles) to the gene selection process induced by typical supervised learning tasks in microarray studies. We originate sample-tracking profiles as the aggregated off-training evaluation of SVM models of increasing gene panel sizes. Genes are ranked by E-RFE, an entropy-based variant of the recursive feature elimination for support vector machines (RFE-SVM). A Dynamic Time Warping (DTW) algorithm is then applied to define a metric between sample-tracking profiles. An unsupervised clustering based on the DTW metric allows automating the discovery of outliers and of subtypes of different molecular profiles. Applications are described on synthetic data and in two gene expression studies.

Algorithms↗

Segmentation methodology for automated classification and differentiation of soft tissues in multiband images of high-resolution ultrasonic transmission tomography.

This paper presents a novel segmentation methodology for automated classification and differentiation of soft tissues using multiband data obtained with the newly developed system of high-resolution ultrasonic transmission tomography (HUTT) for imaging biological organs. This methodology extends and combines two existing approaches: the L-level set active contour (AC) segmentation approach and the agglomerative hierarchical kappa-means approach for unsupervised clustering (UC). To prevent the trapping of the current iterative minimization AC algorithm in a local minimum, we introduce a multiresolution approach that applies the level set functions at successively increasing resolutions of the image data. The resulting AC clusters are subsequently rearranged by the UC algorithm that seeks the optimal set of clusters yielding the minimum within-cluster distances in the feature space. The presented results from Monte Carlo simulations and experimental animal-tissue data demonstrate that the proposed methodology outperforms other existing methods without depending on heuristic parameters and provides a reliable means for soft tissue differentiation in HUTT images.

Algorithms↗

Probabilistic space-time video modeling via piecewise GMM.

In this paper, we describe a statistical video representation and modeling scheme. Video representation schemes are needed to segment a video stream into meaningful video-objects, useful for later indexing and retrieval applications. In the proposed methodology, unsupervised clustering via Gaussian mixture modeling extracts coherent space-time regions in feature space, and corresponding coherent segments (video-regions) in the video content. A key feature of the system is the analysis of video input as a single entity as opposed to a sequence of separate frames. Space and time are treated uniformly. The probabilistic space-time video representation scheme is extended to a piecewise GMM framework in which a succession of GMMs are extracted for the video sequence, instead of a single global model for the entire sequence. The piecewise GMM framework allows for the analysis of extended video sequences and the description of nonlinear, nonconvex motion patterns. The extracted space-time regions allow for the detection and recognition of video events. Results of segmenting video content into static versus dynamic video regions and video content editing are presented.

Algorithms↗

Semantic categorization in the human brain: spatiotemporal dynamics revealed by magnetoencephalography.

We examined the cortical representation of semantic categorization using magnetic source imaging in a task that revealed both dissociations among superordinate categories and associations among different base-level concepts within these categories. Around 200 ms after stimulus onset, the spatiotemporal correlation of brain activity elicited by base-level concepts was greater within than across superordinate categories in the right temporal lobe. Unsupervised clustering of data showed similar categorization between 210 and 450 ms mainly in the left hemisphere. This pattern suggests that well-defined semantic categories are represented in spatially distinct, macroscopically separable neural networks, independent of physical stimulus properties. In contrast, a broader, task-required categorization (natural/man-made) was not evident in our data. The perceptual dynamics of the categorization process is initially evident in the extrastriate areas of the right hemisphere; this activation is followed by higher-level activity along the ventral processing stream, implicating primarily the left temporal lobe.

Adult↗

Transcriptional profiles in melanocytes from clinically unaffected skin distinguish the neoplastic growth pattern in patients with melanoma.

BACKGROUND: It is generally accepted that sunlight may contribute to the development of melanoma. OBJECTIVES: To analyse gene expression of melanocytes obtained from clinically unaffected skin of patients with melanoma and healthy controls before and after exposure to ultraviolet B radiation. METHODS: Using GeneChip array technology, the gene expression of melanocytes obtained from the two donor groups was profiled, in order to identify transcriptional differences affecting susceptibility to melanoma. RESULTS: The data collected did not show any difference between the expression profiles of melanocytes purified from normal donors and from patients with melanoma that was able to give a statistically significant class separation. However, by means of unsupervised clustering our data could be divided into two main classes. The first class included the transcriptome profiles of melanocytes obtained from skin samples of patients with a vertical growth phase (VGP) melanoma, while the second class included the transcriptome profiles of melanocytes obtained from skin samples of patients with a radial growth phase (RGP) melanoma. CONCLUSIONS: These data suggest that melanocytes in patients with VGP and RGP melanomas show significant differences in gene expression profiles, which allow us to classify patients with melanoma also from clinically unaffected skin.

Adult↗

Profiling of apoptosis genes allows for clinical stratification of primary nodal diffuse large B-cell lymphomas.

Intrinsic resistance of lymphoma cells to apoptosis is a probable mechanism causing chemotherapy resistance and eventual fatal outcome in patients with diffuse large B cell lymphomas (DLBCL). We investigated whether microarray expression profiling of apoptosis related genes predicts clinical outcome in 46 patients with primary nodal DLBCL. Unsupervised cluster analysis using genes involved in apoptosis (n = 246) resulted in three separate DLBCL groups partly overlapping with germinal centre B-lymphocytes versus activated B-cells like phenotype. One group with poor clinical outcome was characterised by high expression levels of pro-and anti-apoptotic genes involved in the intrinsic apoptosis pathway. A second group, also with poor clinical outcome, was characterised by high levels of apoptosis inducing cytotoxic effector genes, possibly reflecting a cellular cytotoxic immune response. The third group showing a favourable outcome was characterised by low expression levels of genes characteristic for both other groups. Our results suggest that chemotherapy refractory DLBCL are characterised either by an intense cellular cytotoxic immune response or by constitutive activation of the intrinsic mediated apoptosis pathway with concomitant downstream inhibition of this apoptosis pathway. Consequently, strategies neutralising the function of apoptosis-inhibiting proteins might be effective as alternative treatment modality in part of chemotherapy refractory DLBCL.

Adult↗

Gene expression profiling reveals unique molecular subtypes of Neurofibromatosis Type I-associated and sporadic malignant peripheral nerve sheath tumors.

Malignant peripheral nerve sheath tumors (MPNSTs) are highly aggressive Schwann cell neoplasms that are frequently associated with Type I Neurofibromatosis (NF1) and respond poorly to current therapeutic regimens. To better understand the molecular heterogeneity of these tumors, we performed gene expression profiling on 25 NF1-associated and 17 sporadic MPNSTs using oligonucleotide microarrays representing approximately 8100 unique human gene transcripts. Using several previously reported statistical approaches, we were unable to identify a molecular signature that could reliably distinguish between NF1-associated and sporadic MPNSTs in independent training and test sample sets. However, using an unsupervised clustering approach, we identified an extensive gene expression signature that distinguished 9 of the 42 tumors analyzed. This signature corresponded to relative overexpression of transcripts associated with neuroglial differentiation (NCAM, MBP, L1CAM, P1P) and relative down-regulation of proliferation and growth factor associated transcripts (IGF2, FGFR1, MDK, Ki67). All tumors with this gene expression signature lacked expression of EGFR and all but one tumor were derived from patients with NF1. However, there were no other obvious associations with histological grade, tumor site, metastasis, recurrence, age, or patient survival. We conclude that distinct molecular classes of MPNST exist and that the ability to stratify these tumors based on unique and biologically relevant gene expression profiles may be important for future targeted therapeutics.

Adolescent↗

From quantitative microscopy to automated image understanding.

Quantitative microscopy has been extensively used in biomedical research and has provided significant insights into structure and dynamics at the cell and tissue level. The entire procedure of quantitative microscopy is comprised of specimen preparation, light absorption/reflection/emission from the specimen, microscope optical processing, optical/electrical conversion by a camera or detector, and computational processing of digitized images. Although many of the latest digital signal processing techniques have been successfully applied to compress, restore, and register digital microscope images, automated approaches for recognition and understanding of complex subcellular patterns in light microscope images have been far less widely used. We describe a systematic approach for interpreting protein subcellular distributions using various sets of subcellular location features (SLF), in combination with supervised classification and unsupervised clustering methods. These methods can handle complex patterns in digital microscope images, and the features can be applied for other purposes such as objectively choosing a representative image from a collection and performing statistical comparisons of image sets.

Algorithms↗

Expression of receptor activator of nuclear factor kappabeta ligand (RANKL) and tumour necrosis factor related, apoptosis inducing ligand (TRAIL) in breast cancer, and their relations with osteoprotegerin, oestrogen receptor, and clinicopathological variables.

BACKGROUND: Receptor activator of nuclear factor kappabeta ligand (RANKL) has an important role in bone remodelling, and tumour necrosis factor related, apoptosis inducing ligand (TRAIL) can induce apoptosis in cancer cells. Their functions are linked by their interactions with osteoprotegerin (OPG). OBJECTIVE: To investigate the expression of RANKL and TRAIL in a large series of unselected breast cancers and to analyse the relations between these expressions and the expression of OPG, oestrogen receptor, and clinicopathological variables. METHODS: 395 breast cancers were sampled into tissue microarrays and immunohistochemistry undertaken for RANKL and TRAIL. RESULTS: There was strong expression of RANKL in 14% of the cancers and strong expression of TRAIL in 30%. Expression of RANKL had a negative association with expression of oestrogen receptor (p = 0.036). Expression of TRAIL had a negative association with the Nottingham Prognostic Index (p = 0.021). There was a significant negative relation between expression of RANKL and TRAIL (p<0.005). Unsupervised cluster analysis produced a dendrogram that showed a clear division into two groups, and the expression of oestrogen receptor was significantly higher in one of those groups (p = 0.012). CONCLUSIONS: There is apparent loss of expression of RANKL in 86% of breast cancers; those tumours that retain expression tend to be oestrogen receptor negative and of a high histological grade. There is strong expression of TRAIL in 30% of breast cancers and these tend to be of better prognostic type. These results may be important in the processes of metastasis to bone and the apoptotic cell death pathway in cancer.

Apoptosis Regulatory Proteins↗

Predicting gene function from gene expressions and ontologies.

We introduce a methodology for inducing predictive rule models for functional classification of gene expressions from microarray hybridisation experiments. The basic learning method is the rough set framework for rule induction. The methodology is different from the commonly used unsupervised clustering approaches in that it exploits background knowledge of gene function in a supervised manner. Genes are annotated using Ashburner's Gene Ontology and the functional classes used for learning are mined from these annotations. From the original expression data, we extract a set of biologically meaningful features that are used for learning. A rule model is induced from the data described in terms of these features. Its predictive quality is fine-turned via cross-validation on subsets of the known genes prior to classification of unknown genes. The predictive and descriptive quality of such a rule model is demonstrated on the fibroblast serum response data previously analysed by Iyer et. al. Our analysis shows that the rules are capable of representing the complex relationship between gene expressions and function, and that it is possible to put forward high quality hypotheses about the function of unknown genes.

Algorithms↗

Automatic determination of radial basis functions: an immunity-based approach.

The appropriate operation of a radial basis function (RBF) neural network depends mainly upon an adequate choice of the parameters of its basis functions. The simplest approach to train an RBF network is to assume fixed radial basis functions defining the activation of the hidden units. Once the RBF parameters are fixed, the optimal set of output weights can be determined straightforwardly by using a linear least squares algorithm, which generally means reduction in the learning time as compared to the determination of all RBF network parameters using supervised learning. The main drawback of this strategy is the requirement of an efficient algorithm to determine the number, position, and dispersion of the RBFs. The approach proposed here is inspired by models derived from the vertebrate immune system, that will be shown to perform unsupervised cluster analysis. The algorithm is introduced and its performance is compared to that of the random, k-means center selection procedures and other results from the literature. By automatically defining the number of RBF centers, their positions and dispersions, the proposed method leads to parsimonious solutions. Simulation results are reported concerning regression and classification problems.

Algorithms↗

Electrophysiological classification of somatostatin-positive interneurons in mouse sensorimotor cortex.

Classification of inhibitory interneurons is critical in determining their role in normal information processing and pathophysiological conditions such as epilepsy. Classification schemes have relied on morphological, physiological, biochemical, and molecular criteria; and clear correlations have been demonstrated between firing patterns and cellular markers such as neuropeptides and calcium-binding proteins. This molecular diversity has allowed generation of transgenic mouse strains in which GFP expression is linked to the expression of one of these markers and presumably a single subtype of neuron. In the GIN mouse (EGFP-expressing Inhibitory Neurons), a subpopulation of somatostatin-containing interneurons in the hippocampus and neocortex is labeled with enhanced green fluorescent protein (EGFP). To optimize the use of the GIN mouse, it is critical to know whether the population of somatostatin-EGFP-expressing interneurons is homogeneous. We performed unsupervised cluster analysis on 46 EGFP-expressing interneurons, based on data obtained from whole cell patch-clamp recordings. Cells were classified according to a number of electrophysiological variables related to spontaneous excitatory postsynaptic currents (sEPSCs), firing behavior, and intrinsic membrane properties. EGFP-expressing interneurons were heterogeneous and at least four subgroups could be distinguished. In addition, multiple discriminant analysis was applied to data collected during whole cell recordings to develop an algorithm for predicting the group membership of newly encountered EGFP-expressing interneurons. Our data are consistent with a heterogeneous population of neurons based on electrophysiological properties and indicate that EGFP expression in the GIN mouse is not restricted to a single class of somatostatin-positive interneuron.

Algorithms↗

Integrated array-comparative genomic hybridization and expression array profiles identify clinically relevant molecular subtypes of glioblastoma.

Glioblastoma, the most aggressive primary brain tumor in humans, exhibits a large degree of molecular heterogeneity. Understanding the molecular pathology of a tumor and its linkage to behavior is an important foundation for developing and evaluating approaches to clinical management. Here we integrate array-comparative genomic hybridization and array-based gene expression profiles to identify relationships between DNA copy number aberrations, gene expression alterations, and survival in 34 patients with glioblastoma. Unsupervised clustering on either profile resulted in similar groups of patients, and groups defined by either method were associated with survival. The high concordance between these separate molecular classifications suggested a strong association between alterations on the DNA and RNA levels. We therefore investigated relationships between DNA copy number and gene expression changes. Loss of chromosome 10, a predominant genetic change, was associated not only with changes in the expression of genes located on chromosome 10 but also with genome-wide differences in gene expression. We found that CHI3L1/YKL-40 was significantly associated with both chromosome 10 copy number loss and poorer survival. Immortalized human astrocytes stably transfected with CHI3L1/YKL-40 exhibited changes in gene expression similar to patterns observed in human tumors and conferred radioresistance and increased invasion in vitro. Taken together, the results indicate that integrating DNA and mRNA-based tumor profiles offers the potential for a clinically relevant classification more robust than either method alone and provides a basis for identifying genes important in glioma pathogenesis.

Adipokines↗

Genomic and expression profiling of human spermatocytic seminomas: primary spermatocyte as tumorigenic precursor and DMRT1 as candidate chromosome 9 gene.

Spermatocytic seminomas are solid tumors found solely in the testis of predominantly elderly individuals. We investigated these tumors using a genome-wide analysis for structural and numerical chromosomal changes through conventional karyotyping, spectral karyotyping, and array comparative genomic hybridization using a 32 K genomic tiling-path resolution BAC platform (confirmed by in situ hybridization). Our panel of five spermatocytic seminomas showed a specific pattern of chromosomal imbalances, mainly numerical in nature (range, 3-24 per tumor). Gain of chromosome 9 was the only consistent anomaly, which in one case also involved amplification of the 9p21.3-pter region. Parallel chromosome level expression profiling as well as microarray expression analyses (Affymetrix U133 plus 2.0) was also done. Unsupervised cluster analysis showed that a profile containing transcriptional data on 373 genes (difference of > or = 3.0-fold) is suitable for distinguishing these tumors from seminomas/dysgerminomas. The diagnostic markers SSX2-4 and POU5F1 (OCT3/OCT4), previously identified by us, were among the top discriminatory genes, thereby validating the experimental set-up. In addition, novel discriminatory markers suitable for diagnostic purposes were identified, including Deleted in Azospermia (DAZ). Although the seminomas/dysgerminomas were characterized by expression of stem cell-specific genes (e.g., POU5F1, PROM1/CD133, and ZFP42), spermatocytic seminomas expressed multiple cancer testis antigens, including TSP50 and CTCFL (BORIS), as well as genes known to be expressed specifically during prophase meiosis I (TCFL5, CLGN, and LDHc). This is consistent with different cells of origin, the primordial germ cell and primary spermatocyte, respectively. Based on the region of amplification defined on 9p and the associated expression plus confirmatory immunohistochemistry, DMRT1 (a male-specific transcriptional regulator) was identified as a likely candidate gene for involvement in the development of spermatocytic seminomas.

Biomarkers, Tumor↗

Gene expression profiling on lung cancer outcome prediction: present clinical value and future premise.

DNA microarray has been widely used in cancer research to better predict clinical outcomes and potentially improve patient management. The new approach provides accurate tumor classification and outcome predictions, such as tumor stage, metastatic status, and patient survival, and offers some hope for individualized medicine. However, growing evidence suggests that gene-based prediction is not stable and little is known about the prediction power of gene expression profiles compared with well-known clinical and pathologic predictors. This review summarized up-to-date publications in microarray-based lung cancer clinical outcome prediction and conducted secondary analyses for those with sufficient sample sizes and associated clinical information. Among the most commonly used analytic approaches, unsupervised clustering mainly recaptures tumor histology and provides variable degrees of prediction for tumor stage, lymph node status, or survival. Overall, most studies lack an independent validation. Supervised learning and testing generally offer a better prediction. Noted is that when conventional predictors of age, gender, stage, cell type, and tumor grade are considered collectively, the predictive advantage of the gene expression profiles diminishes. We conclude that outcome prediction from gene expression signatures selected by current analytic approaches can be mostly explained by well-known conventional predictors, particularly histologic subtype and grade of differentiation. A strategy for establishing independent or more accurate signatures is commented.

Cell Differentiation↗