Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Estimation of the tissue composition of the tumour mass in neuroblastoma using segmented CT images.

Neuroblastoma is the most common extra-cranial, solid, malignant tumour in children. Advances in radiology have made possible the detection and staging of the disease. Nevertheless, there is no method available at present that can go beyond detection and qualitative analysis, towards quantitative assessment of the tissue composition of the primary tumour mass in neuroblastoma. Such quantitative analysis could provide important information and serve as a decision-support tool to the radiologist and the oncologist, result in better treatment and follow-up and even lead to the avoidance of delayed surgery. The problem investigated was the improvement of the analysis of the primary tumour mass, in patients with neuroblastoma, using X-ray computed tomography (CT) images. A methodology was proposed for the estimation of the tissue content of the mass: it comprised a Gaussian mixture model for estimation, from segmented CT images, of the tissue composition of the primary tumour. To demonstrate the potential of the method, the results are presented of its application to ten CT examinations of four patients. The method provides quantitative information, and it was observed that the tumour in one of the patients reduced from 523 cm3 to 81 cm3 in volume, with an increase in calcification from about 20% to about 88% of the tumour volume, in response to chemotherapy over a period of five months. Results indicate that the proposed technique may be of considerable value in assessing the response to therapy of patients with neuroblastoma.

Child↗

Generative topographic mapping applied to clustering and visualization of motor unit action potentials.

The identification and visualization of clusters formed by motor unit action potentials (MUAPs) is an essential step in investigations seeking to explain the control of the neuromuscular system. This work introduces the generative topographic mapping (GTM), a novel machine learning tool, for clustering of MUAPs, and also it extends the GTM technique to provide a way of visualizing MUAPs. The performance of GTM was compared to that of three other clustering methods: the self-organizing map (SOM), a Gaussian mixture model (GMM), and the neural-gas network (NGN). The results, based on the study of experimental MUAPs, showed that the rate of success of both GTM and SOM outperformed that of GMM and NGN, and also that GTM may in practice be used as a principled alternative to the SOM in the study of MUAPs. A visualization tool, which we called GTM grid, was devised for visualization of MUAPs lying in a high-dimensional space. The visualization provided by the GTM grid was compared to that obtained from principal component analysis (PCA).

Action Potentials↗

Investigations of dipole localization accuracy in MEG using the bootstrap.

We describe the use of the nonparametric bootstrap to investigate the accuracy of current dipole localization from magnetoencephalography (MEG) studies of event-related neural activity. The bootstrap is well suited to the analysis of event-related MEG data since the experiments are repeated tens or even hundreds of times and averaged to achieve acceptable signal-to-noise ratios (SNRs). The set of repetitions or epochs can be viewed as a set of independent realizations of the brain's response to the experiment. Bootstrap resamples can be generated by sampling with replacement from these epochs and averaging. In this study, we applied the bootstrap resampling technique to MEG data from somatotopic experimental and simulated data. Four fingers of the right and left hand of a healthy subject were electrically stimulated, and about 400 trials per stimulation were recorded and averaged in order to measure the somatotopic mapping of the fingers in the S1 area of the brain. Based on single-trial recordings for each finger we performed 5000 bootstrap resamples. We reconstructed dipoles from these resampled averages using the Recursively Applied and Projected (RAP)-MUSIC source localization algorithm. We also performed a simulation for two dipolar sources with overlapping time courses embedded in realistic background brain activity generated using the prestimulus segments of the somatotopic data. To find correspondences between multiple sources in each bootstrap, sample dipoles with similar time series and forward fields were assumed to represent the same source. These dipoles were then clustered by a Gaussian Mixture Model (GMM) clustering algorithm using their combined normalized time series and topographies as feature vectors. The mean and standard deviation of the dipole position and the dipole time series in each cluster were computed to provide estimates of the accuracy of the reconstructed source locations and time series.

Brain Mapping↗

Natural segmentation of the locomotor behavior of drug-induced rats in a photobeam cage.

Recently, Drai et al. (J Neurosci Methods 96 (2000) 119) have introduced an algorithm that segments rodent locomotor behavior into natural units of 'staying in place' (lingering) behavior versus going between places (progression segments). This categorization, based on the maximum speed attained within the segment, was shown to be intrinsic to the data, using the statistical method of Gaussian Mixture Model. These results were obtained in normal rats and mice using very large (650 or 320 cm) circular arenas and a video tracking system. In the present study, we reproduce these results with amphetamine, phencyclidine and saline injected rats, using data measured by a standard photobeam tracking system in square 45 cm cages. An intrinsic distinction between two or three 'gears' could be shown in all animals. The spatial distribution of these gears indicates that, as in the large arena behavior, they correspond to the difference between 'staying in place' behavior and 'going between places'. The robustness of this segmentation over arena size, different measurement system and dose of two psychostimulant drugs indicates that this is an intrinsic, natural segmentation of rodent locomotor behavior. Analysis of photobeam data that is based on this segmentation has thus a potential use in psychopharmacology research.

Algorithms↗

A pseudolikelihood approach for simultaneous analysis of array comparative genomic hybridizations.

DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based comparative genomic hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure and with random effects to allow for intertumoral variation, as well as intratumoral clonal variation. For ease of computation, we base estimation on a pseudolikelihood function. The method produces quantitative assessments of the likelihood of genetic alterations at each clone, along with a graphical display for simple visual interpretation. We assess the characteristics of the method through simulation studies and analysis of a brain tumor aCGH data set. We show that the pseudolikelihood approach is superior to existing methods both in detecting small regions of copy number alteration and in accurately classifying regions of change when intratumoral clonal variation is present. Software for this approach is available at http://www.biostat.harvard.edu/ approximately betensky/papers.html.

Algorithms↗

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals↗

Quantitative trait associated microarray gene expression data analysis.

Selection on phenotypes may cause genetic change. To understand the relationship between phenotype and gene expression from an evolutionary viewpoint, it is important to study the concordance between gene expression and profiles of phenotypes. In this study, we use a novel method of clustering to identify genes whose expression profiles are related to a quantitative phenotype. Cluster analysis of gene expression data aims at classifying genes into several different groups based on the similarity of their expression profiles across multiple conditions. The hope is that genes that are classified into the same clusters may share underlying regulatory elements or may be a part of the same metabolic pathways. Current methods for examining the association between phenotype and gene expression are limited to linear association measured by the correlation between individual gene expression values and phenotype. Genes may be associated with the phenotype in a nonlinear fashion. In addition, groups of genes that share a particular pattern in their relationship to phenotype may be of evolutionary interest. In this study, we develop a method to group genes based on orthogonal polynomials under a multivariate Gaussian mixture model. The effect of each expressed gene on the phenotype is partitioned into a cluster mean and a random deviation from the mean. Genes can also be clustered based on a time series. Parameters are estimated using the expectation-maximization algorithm and implemented in SAS. The method is verified with simulated data and demonstrated with experimental data from 2 studies, one clusters with respect to severity of disease in Alzheimer's patients and another clusters data for a rat fracture healing study over time. We find significant evidence of nonlinear associations in both studies and successfully describe these patterns with our method. We give detailed instructions and provide a working program that allows others to directly implement this method in their own analyses.

Animals↗

A multiscale expectation-maximization semisupervised classifier suitable for badly posed image classification.

This paper deals with the problem of badly posed image classification. Although underestimated in practice, bad-posedness is likely to affect many real-world image classification tasks, where reference samples are difficult to collect (e.g., in remote sensing (RS) image mapping) and/or spatial autocorrelation is relevant. In an image classification context affected by a lack of reference samples, an original inductive learning multiscale image classifier, termed multiscale semisupervised expectation maximization (MSEM), is proposed. The rationale behind MSEM is to combine useful complementary properties of two alternative data mapping procedures recently published outside of image processing literature, namely, the multiscale modified Pappas adaptive clustering (MPAC) algorithm and the sample-based semisupervised expectation maximization (SEM) classifier. To demonstrate its potential utility, MSEM is compared against nonstandard classifiers, such as MPAC, SEM and the single-scale contextual SEM (CSEM) classifier, besides against well-known standard classifiers in two RS image classification problems featuring few reference samples and modestly useful texture information. These experiments yield weak (subjective) but numerous quantitative map quality indexes that are consistent with both theoretical considerations and qualitative evaluations by expert photointerpreters. According to these quantitative results, MSEM is competitive in terms of overall image mapping performance at the cost of a computational overhead three to six times superior to that of its most interesting rival, SEM. More in general, our experiments confirm that, even if they rely on heavy class-conditional normal distribution assumptions that may not be true in many real-world problems (e.g., in highly textured images), semisupervised classifiers based on the iterative expectation maximization Gaussian mixture model solution can be very powerful in practice when: 1) there is a lack of reference samples with respect to the problem/model complexity and 2) texture information is considered negligible (i.e., a piecewise constant image model holds).

Algorithms↗

On classification with incomplete data.

We address the incomplete-data problem in which feature vectors to be classified are missing data (features). A (supervised) logistic regression algorithm for the classification of incomplete data is developed. Single or multiple imputation for the missing data is avoided by performing analytic integration with an estimated conditional density function (conditioned on the observed data). Conditional density functions are estimated using a Gaussian mixture model (GMM), with parameter estimation performed using both Expectation-Maximization (EM) and Variational Bayesian EM (VB-EM). The proposed supervised algorithm is then extended to the semisupervised case by incorporating graph-based regularization. The semisupervised algorithm utilizes all available data-both incomplete and complete, as well as labeled and unlabeled. Experimental results of the proposed classification algorithms are shown.

Algorithms↗

A comparative study on kernel-based probabilistic neural networks for speaker verification.

This paper compares kernel-based probabilistic neural networks for speaker verification based on 138 speakers of the YOHO corpus. Experimental evaluations using probabilistic decision-based neural networks (PDBNNs), Gaussian mixture models (GMMs) and elliptical basis function networks (EBFNs) as speaker models were conducted. The original training algorithm of PDBNNs was also modified to make PDBNNs appropriate for speaker verification. Results show that the equal error rate obtained by PDBNNs and GMMs is less than that of EBFNs (0.33% vs. 0.48%), suggesting that GMM- and PDBNN-based speaker models outperform the EBFN ones. This work also finds that the globally supervised learning of PDBNNs is able to find decision thresholds that not only maintain the false acceptance rates to a low level but also reduce their variation, whereas the ad-hoc threshold-determination approach used by the EBFNs and GMMs causes a large variation in the error rates. This property makes the performance of PDBNN-based systems more predictable.

Algorithms↗

An alternative perspective on adaptive independent component analysis algorithms

This article develops an extended independent component analysis algorithm for mixtures of arbitrary subgaussian and supergaussian sources. The gaussian mixture model of Pearson is employed in deriving a closed-form generic score function for strictly subgaussian sources. This is combined with the score function for a unimodal supergaussian density to provide a computationally simple yet powerful algorithm for performing independent component analysis on arbitrary mixtures of nongaussian sources.

Journal Article↗

Identification of expressed genes linked to malignancy of human colorectal carcinoma by parametric clustering of quantitative expression data.

BACKGROUND: Individual human carcinomas have distinct biological and clinical properties: gene-expression profiling is expected to unveil the underlying molecular features. Particular interest has been focused on potential diagnostic and therapeutic applications. Solid tumors, such as colorectal carcinoma, present additional obstacles for experimental and data analysis. RESULTS: We analyzed the expression levels of 1,536 genes in 100 colorectal cancer and 11 normal tissues using adaptor-tagged competitive PCR, a high-throughput reverse transcription-PCR technique. A parametric clustering method using the Gaussian mixture model and the Bayes inference revealed three groups of expressed genes. Two contained large numbers of genes. One of these groups correlated well with both the differences between tumor and normal tissues and the presence or absence of distant metastasis, whereas the other correlated only with the tumor/normal difference. The third group comprised a small number of genes. Approximately half showed an identical expression pattern, and cancer tissues were classified into two groups by their expression levels. The high-expression group had strong correlation with distant metastasis, and a poorer survival rate than the low-expression group, indicating possible clinical applications of these genes. In addition to c-yes, a homolog of a viral oncogene, prognostic indicators included genes specific to glial cells, which gives a new link between malignancy and ectopic gene expression. CONCLUSIONS: The malignancy of human colorectal carcinoma is correlated with a unique expression pattern of a specific group of genes, allowing the classification of tumor tissues into two clinically distinct groups.

Cluster Analysis↗

Exploration of statistical dependence between illness parameters using the entropy correlation coefficient.

UNLABELLED: The entropy correlation coefficient (ECC) is a useful tool for measuring statistical dependence between variables. We employed this tool to search for pairs of variables that correlated in the chronic fatigue syndrome (CFS) Computational Challenge dataset. Highly related variables are candidates for data reduction, and novel relationships could lead to hypotheses regarding the pathogenesis of CFS. METHODS: Data for 130 female participants in the Wichita (KS, USA) clinical study [1] was coded into numerical values. Metric data was grouped using Gaussian mixture models; the number of groups was chosen using Bayesian information content. The pair-wise correlation between all variables was computed using the ECC. Significance was estimated from 1000 iterations of a permutation test and a threshold of 0.01 was used to identify significantly correlated variables. RESULTS: The five dimensions of multidimensional fatigue inventory (MFI) were all highly correlated with each other. Seven Short Form (SF)-36 measures, four CFS case-defining symptoms and the Zung self-rating depression scale all correlated with all MFI dimensions. No physiological variables correlate with more than one MFI dimension. MFI, SF-36, CDC symptom inventory, the Zung self-rating depression scale and three Cambridge Neuropsychological Test Automated Battery (CANTAB) measures are highly correlated with CFS disease status. DISCUSSION: Correlations between the five dimensions of MFI are expected since they are measured from the same instrument. The relationship between MFI and Zung depression index has been previously reported. MFI, SF-36, and Centers for Disease Control and Prevention (CDC) symptom inventory are used to classify CFS; it is not surprising that they are correlated with disease status. Only one of the three CANTAB measures that correlate with disease status has been previously found, indicating the ECC identifies relationships not found with other statistical tools. CONCLUSION: The ECC is a useful tool for measuring statistical dependence between variables in clinical and laboratory datasets. The ECC needs to be further studied to gain a better understanding of its meaning for clinical data.

Adult↗

A feasibility study of multispectral image analysis of skin tumors.

To develop a noninvasive, early-detection method for skin cancers, a feasibility study of multispectral image analysis was investigated. The three most frequently occurring skin cancer types, ten basal-cell carcinomas (BCCs), ten squamous-cell carcinomas (SCCs) and five malignant melanomas (MMs) were studied, along with ten normal moles. Images were acquired by a charge-coupled device camera using eight narrow-band filters ranging from 450 nm to 800 nm, at 50-nm intervals. To extract main features of these tumors, principal components analysis (PCA) was performed, because it projects the multidimensional (here, eight-dimensional) data in the direction of maximum data variance. Then, the primary PCA components for red, green, and blue subset images were analyzed in terms of hue-saturation-intensity (HSI). By hue distributions, the BCCs and SCCs were differentiated from the MMs and normal moles. Texture information was used to further classify tumor types after the HSI analysis. The texture analysis, performed using a spatial gray-level co-occurrence matrix (SGCM), could separate MMs from normal moles. The BCCs and SCCs were further studied by Fisher's linear discriminant analysis. Distribution was described as a Gaussian mixture model. By this classification procedure, seven BCCs, eight SCCs, five MMs, and ten NMs were correctly classified. Three BCCs and two SCCs were unseparable. Thus, multispectral skin cancer image analysis has potential to diagnose skin cancers.

Algorithms↗

A dissimilarity matrix between protein atom classes based on Gaussian mixtures.

MOTIVATION: Previously, Rantanen et al. (2001; J. Mol. Biol., 313, 197-214) constructed a protein atom-ligand fragment interaction library embodying experimentally solved, high-resolution three-dimensional (3D) structural data from the Protein Data Bank (PDB). The spatial locations of protein atoms that surround ligand fragments were modeled with Gaussian mixture models, the parameters of which were estimated with the expectation-maximization (EM) algorithm. In the validation analysis of this library, there was strong indication that the protein atom classification, 24 classes, was too large and that a reduction in the classes would lead to improved predictions. RESULTS: Here, a dissimilarity (distance) matrix that is suitable for comparison and fusion of 24 pre-defined protein atom classes has been derived. Jeffreys' distances between Gaussian mixture models are used as a basis to estimate dissimilarities between protein atom classes. The dissimilarity data are analyzed both with a hierarchical clustering method and independently by using multidimensional scaling analysis. The results provide additional insight into the relationships between different protein atom classes, giving us guidance on, for example, how to readjust protein atom classification and, thus, they will help us to improve protein--ligand interaction predictions. CONTACT: vira@utu.fi

Cluster Analysis↗

Soft mixer assignment in a hierarchical generative model of natural scene statistics.

Gaussian scale mixture models offer a top-down description of signal generation that captures key bottom-up statistical characteristics of filter responses to images. However, the pattern of dependence among the filters for this class of models is prespecified. We propose a novel extension to the gaussian scale mixture model that learns the pattern of dependence from observed inputs and thereby induces a hierarchical representation of these inputs. Specifically, we propose that inputs are generated by gaussian variables (modeling local filter structure), multiplied by a mixer variable that is assigned probabilistically to each input from a set of possible mixers. We demonstrate inference of both components of the generative model, for synthesized data and for different classes of natural images, such as a generic ensemble and faces. For natural images, the mixer variable assignments show invariances resembling those of complex cells in visual cortex; the statistics of the gaussian components of the model are in accord with the outputs of divisive normalization models. We also show how our model helps interrelate a wide range of models of image statistics and cortical processing.

Animals↗

A Bayesian molecular interaction library.

We describe a library of molecular fragments designed to model and predict non-bonded interactions between atoms. We apply the Bayesian approach, whereby prior knowledge and uncertainty of the mathematical model are incorporated into the estimated model and its parameters. The molecular interaction data are strengthened by narrowing the atom classification to 14 atom types, focusing on independent molecular contacts that lie within a short cutoff distance, and symmetrizing the interaction data for the molecular fragments. Furthermore, the location of atoms in contact with a molecular fragment are modeled by Gaussian mixture densities whose maximum a posteriori estimates are obtained by applying a version of the expectation-maximization algorithm that incorporates hyperparameters for the components of the Gaussian mixtures. A routine is introduced providing the hyperparameters and the initial values of the parameters of the Gaussian mixture densities. A model selection criterion, based on the concept of a 'minimum message length' is used to automatically select the optimal complexity of a mixture model and the most suitable orientation of a reference frame for a fragment in a coordinate system. The type of atom interacting with a molecular fragment is predicted by values of the posterior probability function and the accuracy of these predictions is evaluated by comparing the predicted atom type with the actual atom type seen in crystal structures. The fact that an atom will simultaneously interact with several molecular fragments forming a cohesive network of interactions is exploited by introducing two strategies that combine the predictions of atom types given by multiple fragments. The accuracy of these combined predictions is compared with those based on an individual fragment. Exhaustive validation analyses and qualitative examples (e.g., the ligand-binding domain of glutamate receptors) demonstrate that these improvements lead to effective modeling and prediction of molecular interactions.

Algorithms↗

BYY harmony learning, structural RPCL, and topological self-organizing on mixture models.

The Bayesian Ying-Yang (BYY) harmony learning acts as a general statistical learning framework, featured by not only new regularization techniques for parameter learning but also a new mechanism that implements model selection either automatically during parameter learning or via a new class of model selection criteria used after parameter learning. In this paper, further advances on BYY harmony learning by considering modular inner representations are presented in three parts. One consists of results on unsupervisedmixture models, ranging from Gaussian mixture based Mean Square Error (MSE) clustering, elliptic clustering, subspace clustering to NonGaussian mixture based clustering not only with each cluster represented via either Bernoulli-Gaussian mixtures or independent real factor models, but also with independent component analysis implicitly made on each cluster. The second consists of results on supervised mixture-of-experts (ME) models, including Gaussian ME, Radial Basis Function nets, and Kernel regressions. The third consists of two strategies for extending the above structural mixtures into self-organized topological maps. All these advances are introduced with details on three issues, namely, (a) adaptive learning algorithms, especially elliptic, subspace, and structural rival penalized competitive learning algorithms, with model selection made automatically during learning; (b) model selection criteria for being used after parameter learning, and (c) how these learning algorithms and criteria are obtained from typical special cases of BYY harmony learning.

Bayes Theorem↗