Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Segmentation of the fibro-glandular disc in mammograms using Gaussian mixture modelling.

The paper presents a technique for the segmentation of the fibro-glandular disc in mammograms based upon a statistical model of breast density. The density function of the model was represented by a mixture of up to four weighted Gaussians, each one corresponding to a specific density class in the breast. The parameters of the model and the number of tissue classes in the breast were determined using the expectation-maximisation algorithm and the minimum description length method. Grey-level statistics of the pectoral muscle were used to determine the tissue categories that are likely to represent the fibro-glandular disc. The method was applied to 84 medio-lateral oblique mammograms from the Mini-MIAS database. The results of the segmented fibro-glandular disc were assessed by a radiologist using the original and the segmented images, with reference to a ranking table categorising the results of segmentation as: 1: excellent; 2: good; 3: average; 4: poor; and 5: complete failure. Of the 84 cases analysed, 64.3% were rated as excellent, 16.7% were rated as good, 10.7% were rated as average, and 4.7% were rated as poor; only 3.6% of the cases were rated as a complete failure with regard to segmentation of the fibro-glandular disc.

Breast Neoplasms↗

Prediction of random effects in finite mixture models with Gaussian components.

Prediction of random effects in finite mixture models with Gaussian distributions is discussed from a non-Bayesian perspective, assuming that location and dispersion parameters are known. The focus is on calculating the best predictor, that is, the statistic with minimum expected squared prediction error, for several models. Coverage includes mixture sampling models, as well as mixtures for the distribution of the random effects. Longitudinal and cross-sectional specifications with correlated random effects, such as those arising in animal breeding and genetics, are examined. The best linear predictor and the best linear unbiased predictor are derived for these models as well.

Analysis of Variance↗

A spatially constrained mixture model for image segmentation.

Gaussian mixture models (GMMs) constitute a well-known type of probabilistic neural networks. One of their many successful applications is in image segmentation, where spatially constrained mixture models have been trained using the expectation-maximization (EM) framework. In this letter, we elaborate on this method and propose a new methodology for the M-step of the EM algorithm that is based on a novel constrained optimization formulation. Numerical experiments using simulated images illustrate the superior performance of our method in terms of the attained maximum value of the objective function and segmentation accuracy compared to previous implementations of this approach.

Neural Networks, Computer↗

Nonlinear prediction for Gaussian mixture image models.

Prediction is an essential operation in many image processing applications, such as object detection and image and video compression. When the images are modeled as Gaussian, the optimal predictor is linear and easy to obtain. However, image texture and clutter are often non-Gaussian, and, in such cases, optimal predictors are difficult to obtain. In this paper, we derive an optimal predictor for an important class of non-Gaussian image models, the block-based multivariate Gaussian mixture model. This predictor has a special nonlinear structure: it is a linear combination of the neighboring pixels, but the combination coefficients are also functions of the neighboring pixels, not constants. The efficacy of this predictor is demonstrated in object detection experiments where the prediction error image is used to identify "hidden" objects. Experimental results indicate that when the background texture is nonlinear, i.e., with fast-switching gray-level patches, it performs significantly better than the optimal linear predictor.

Algorithms↗

Robust Bayesian clustering.

A new variational Bayesian learning algorithm for Student-t mixture models is introduced. This algorithm leads to (i) robust density estimation, (ii) robust clustering and (iii) robust automatic model selection. Gaussian mixture models are learning machines which are based on a divide-and-conquer approach. They are commonly used for density estimation and clustering tasks, but are sensitive to outliers. The Student-t distribution has heavier tails than the Gaussian distribution and is therefore less sensitive to any departure of the empirical distribution from Gaussianity. As a consequence, the Student-t distribution is suitable for constructing robust mixture models. In this work, we formalize the Bayesian Student-t mixture model as a latent variable model in a different way from Svensén and Bishop [Svensén, M., & Bishop, C. M. (2005). Robust Bayesian mixture modelling. Neurocomputing, 64, 235-252]. The main difference resides in the fact that it is not necessary to assume a factorized approximation of the posterior distribution on the latent indicator variables and the latent scale variables in order to obtain a tractable solution. Not neglecting the correlations between these unobserved random variables leads to a Bayesian model having an increased robustness. Furthermore, it is expected that the lower bound on the log-evidence is tighter. Based on this bound, the model complexity, i.e. the number of components in the mixture, can be inferred with a higher confidence.

Algorithms↗

Acoustic detection and classification of Microchiroptera using machine learning: lessons learned from automatic speech recognition.

Current automatic acoustic detection and classification of microchiroptera utilize global features of individual calls (i.e., duration, bandwidth, frequency extrema), an approach that stems from expert knowledge of call sonograms. This approach parallels the acoustic phonetic paradigm of human automatic speech recognition (ASR), which relied on expert knowledge to account for variations in canonical linguistic units. ASR research eventually shifted from acoustic phonetics to machine learning, primarily because of the superior ability of machine learning to account for signal variation. To compare machine learning with conventional methods of detection and classification, nearly 3000 search-phase calls were hand labeled from recordings of five species: Pipistrellus bodenheimeri, Molossus molossus, Lasiurus borealis, L. cinereus semotus, and Tadarida brasiliensis. The hand labels were used to train two machine learning models: a Gaussian mixture model (GMM) for detection and classification and a hidden Markov model (HMM) for classification. The GMM detector produced 4% error compared to 32% error for a baseline broadband energy detector, while the GMM and HMM classifiers produced errors of 0.6 +/- 0.2% compared to 16.9 +/- 1.1% error for a baseline discriminant function analysis classifier. The experiments showed that machine learning algorithms produced errors an order of magnitude smaller than those for conventional methods.

Acoustics↗

Edgeworth-expanded gaussian mixture density modeling.

Instead of increasing the order of the Edgeworth expansion of a single gaussian kernel, we suggest using mixtures of Edgeworth-expanded gaussian kernels of moderate order. We introduce a simple closed-form solution for estimating the kernel parameters based on weighted moment matching. Furthermore, we formulate the extension to the multivariate case, which is not always feasible with algebraic density approximation procedures.

Algorithms↗

Hidden Markov models for wavelet-based blind source separation.

In this paper, we consider the problem of blind source separation in the wavelet domain. We propose a Bayesian estimation framework for the problem where different models of the wavelet coefficients are considered: the independent Gaussian mixture model, the hidden Markov tree model, and the contextual hidden Markov field model. For each of the three models, we give expressions of the posterior laws and propose appropriate Markov chain Monte Carlo algorithms in order to perform unsupervised joint blind separation of the sources and estimation of the mixing matrix and hyper parameters of the problem. Indeed, in order to achieve an efficient joint separation and denoising procedures in the case of high noise level in the data, a slight modification of the exposed models is presented: the Bernoulli-Gaussian mixture model, which is equivalent to a hard thresholding rule in denoising problems. A number of simulations are presented in order to highlight the performances of the aforementioned approach: 1) in both high and low signal-to-noise ratios and 2) comparing the results with respect to the choice of the wavelet basis decomposition.

Algorithms↗

Scalable model-based clustering for large databases based on data summarization.

The scalability problem in data mining involves the development of methods for handling large databases with limited computational resources such as memory and computation time. In this paper, two scalable clustering algorithms, bEMADS and gEMADS, are presented based on the Gaussian mixture model. Both summarize data into subclusters and then generate Gaussian mixtures from their data summaries. Their core algorithm, EMADS, is defined on data summaries and approximates the aggregate behavior of each subcluster of data under the Gaussian mixture model. EMADS is provably convergent. Experimental results substantiate that both algorithms can run several orders of magnitude faster than expectation-maximization with little loss of accuracy.

Algorithms↗

In silico phenotypic screening method of mutants based on statistical modeling of genetically mixed samples.

In comprehensive functional genomics projects, systematic analysis of phenotypes is vital. However, conventional phenotypic screening is done mainly by imprecise visual observation of qualitative traits, and, therefore, in silico screening techniques for quantitative traits are required. In this report, we propose in silico phenotypic screening method that utilizes a Gaussian mixture model for the trait distribution in the offspring of a mutagenized line and the likelihood ratio test between the estimated Gaussian mixture model and the wild-type single Gaussian model. In order to evaluate the proposed method, we performed a screening experiment using real trait data of Arabidopsis. In this experiment, the proposed screening method properly distinguished the mutant line from the wild-type line. Furthermore, we conducted power analysis of the proposed method and two conventional methods under various simulated conditions of sample size and distribution of trait frequency. The result of the power analysis confirmed the effectiveness of the proposed method compared to the conventional methods.

Algorithms↗

Classification of births by birth weight and gestational age: an application of multivariate mixture models.

BACKGROUND: Classifications such as low birth weight, premature, and small for gestational age. i.e. compromised births, have been criticized because they depend upon arbitrary standards that may not be appropriate for all populations. AIM: This study applies multivariate Gaussian mixture models with covariates to characterize birth weight by gestational age distributions. SUBJECTS AND METHODS: The data consist of Asian, African, Hispanic and European American births in New York State in 1988. The analysis employs maximum likelihood methods. RESULTS: Birth cohorts appear heterogeneous and composed of at least two sub-populations. One sub-population accounts for the majority of births, has a higher mean birth weight and gestational age but small variances. The other sub-population has a lower mean birthweight and gestational age but very large variances. As a result of the large variances this sub-population accounts for compromised births. The model also suggests that a number of compromised births occur within the normal birth weight and gestational age range. Among normal births, birth weight increases and gestational age declines with maternal age. The effects on compromised births vary among populations. CONCLUSIONS: Multivariate Gaussian mixture models provide a method of identifying compromised births that is not dependent upon arbitrary standards.

Adult↗

Prediction of drug release profiles using an intelligent learning system: an experimental study in transdermal iontophoresis.

This paper investigates the use of a neural-network-based intelligent learning system for the prediction of drug release profiles. An experimental study in transdermal iontophoresis (TI) is employed to evaluate the applicability of a particular neural network (NN) model, i.e. the Gaussian mixture model (GMM), in modeling and predicting drug release profiles. A number of tests are systematically designed using the face-centered central composite design (CCD) approach to examine the effects of various process variables simultaneously during the iontophoresis process. The GMM is then applied to model and predict the drug release profiles based on the data samples collected from the experiments. The GMM results are compared with those from multiple regression models. In addition, the bootstrap method is used to assess the reliability of the network predictions by estimating confidence intervals associated with the results. The results demonstrate that the combination of the face-centered CCD and GMM can be employed as a useful intelligent tool for the prediction of time-series profiles in pharmaceutical and biomedical experiments.

Administration, Cutaneous↗

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans↗

Identification of degenerate neuronal systems based on intersubject variability.

Group studies implicitly assume that all subjects activate one common system to sustain a particular cognitive task. Intersubject variability is generally treated as well-behaved and uninteresting noise. However, intersubject variability might result from subjects engaging different degenerate neuronal systems that are each sufficient for task performance. This would produce a multimodal distribution of intersubject variability. We have explored this idea with the help of Gaussian Mixture Modeling and Bayesian model comparison procedures. We illustrate our approach using a crossmodal priming paradigm, in which subjects perform a semantic decision on environmental sounds or their spoken names that were preceded by a semantically congruent or incongruent picture or written name. All subjects consistently activated the superior temporal gyri bilaterally, the left fusiform gyrus and the inferior frontal sulcus. Comparing a One and Two Gaussian Mixture Model of the unexplained residuals provided very strong evidence for two groups with distinct activation patterns: 6 subjects exhibited additional activations in the superior temporal sulci bilaterally, the right superior frontal and central sulcus. 11 subjects showed increased activation in the striate and the right inferior parietal cortex. These results suggest that semantic decisions on auditory-visual compound stimuli might be accomplished by two overlapping degenerate neuronal systems.

Acoustic Stimulation↗

Identification of gastroenteric viruses by electron microscopy using higher order spectral features.

BACKGROUND: Many paediatric illnesses are caused by viral agents, for example, acute gastroenteritis. Electron microscopy can provide images of viral particles and can be used to identify the agents. OBJECTIVES: The use of electron microscopy as a diagnostic tool is limited by the need for high level of expertise in interpreting these images and the time required. A semi-automated method is proposed in this paper. STUDY DESIGN: The method is based on bispectal features that capture contour and texture information while providing robustness to shift, rotation, changes in size and noise. The magnification or true size of the viral particles need not be known precisely, but if available can be used additionally for improved classification. Viral particles from one or more images are segmented and analyzed to verify whether they belong to a particular class (such as Adenovirus, Rotavirus, etc.) or not. Two experiments were conducted-depending on the populations from which virus particle images were collected for training and testing, respectively. In the first, disjoint subsets from a pooled population of virus particles obtained from several images were used. In the second, separate populations from separate images were used. The performance of the method on viruses of similar size was separately evaluated using Astrovirus, HAV and Poliovirus. A Gaussian Mixture Model was used for the probability density of the features. A threshold on the log-likelihood is varied to study false alarm and false rejection trade-off. Features from many particles and/or likelihoods from independent tests are averaged to yield better performance. RESULTS: An equal error rate (EER) of 2% is obtained for verification of Rotavirus (tested against three other viruses) when features from 15 viral particle images are averaged. It drops further to less than 0.2% when scores from two tests are averaged to make a decision. For verification of Astrovirus (tested against two others of the same size) the EER was less than 2% when 20 particles and two tests were used. CONCLUSION: Bispectral features and Gaussian mixture modelling of their probability density are shown to be effective in identifying viruses from electron microscope images. With the use of digital imaging in electron microscopes, this method can be fully automated.

Adenoviruses, Human↗

Model-based clustering and data transformations for gene expression data.

MOTIVATION: Clustering is a useful exploratory technique for the analysis of gene expression data. Many different heuristic clustering algorithms have been proposed in this context. Clustering algorithms based on probability models offer a principled alternative to heuristic algorithms. In particular, model-based clustering assumes that the data is generated by a finite mixture of underlying probability distributions such as multivariate normal distributions. The issues of selecting a 'good' clustering method and determining the 'correct' number of clusters are reduced to model selection problems in the probability framework. Gaussian mixture models have been shown to be a powerful tool for clustering in many applications. RESULTS: We benchmarked the performance of model-based clustering on several synthetic and real gene expression data sets for which external evaluation criteria were available. The model-based approach has superior performance on our synthetic data sets, consistently selecting the correct model and the number of clusters. On real expression data, the model-based approach produced clusters of quality comparable to a leading heuristic clustering algorithm, but with the key advantage of suggesting the number of clusters and an appropriate model. We also explored the validity of the Gaussian mixture assumption on different transformations of real data. We also assessed the degree to which these real gene expression data sets fit multivariate Gaussian distributions both before and after subjecting them to commonly used data transformations. Suitably chosen transformations seem to result in reasonable fits. AVAILABILITY: MCLUST is available at http://www.stat.washington.edu/fraley/mclust. The software for the diagonal model is under development. CONTACT: kayee@cs.washington.edu. SUPPLEMENTARY INFORMATION: http://www.cs.washington.edu/homes/kayee/model.

Algorithms↗

Density estimation by mixture models with smoothing priors

In the statistical approach for self-organizing maps (SOMs), learning is regarded as an estimation algorithm for a gaussian mixture model with a gaussian smoothing prior on the centroid parameters. The values of the hyperparameters and the topological structure are selected on the basis of a statistical principle. However, since the component selection probabilities are fixed to a common value, the centroids concentrate on areas with high data density. This deforms a coordinate system on an extracted manifold and makes smoothness evaluation for the manifold inaccurate. In this article, we study an extended SOM model whose component selection probabilities are variable. To stabilize the estimation, a smoothing prior on the component selection probabilities is introduced. An estimation algorithm for the parameters and the hyperparameters based on empirical Bayesian inference is obtained. The performance of density estimation by the new model and the SOM model is compared via simulation experiments.

Journal Article↗

Dynamic On-line Clustering and State Extraction: An Approach to Symbolic Learning.

Although recurrent neural nets have been moderately successful in learning to emulate finite-state machines (FSMs), the continuous internal state dynamics of a neural net are not well matched to the discrete behavior of an FSM. We describe an architecture, called DOLCE, that allows discrete states to evolve in a net as learning progresses. DOLCE consists of a standard recurrent neural net trained by gradient descent and an adaptive clustering technique that quantizes the state space. We describe two implementations of DOLCE. The first implementation, called DOLCE(u), uses an adaptive clustering scheme in an unsupervised mode to determine both the number of clusters and the partitioning of the state space as learning progresses. The second model, DOLCE(s), uses a Gaussian Mixture Model in a supervised learning framework to infer the states of an FSM. DOLCE(s) is based on the assumption that a finite set of discrete internal states is required for the task, and that the actual network state belongs to this set but has been corrupted by noise due to inaccuracy in the weights. DOLCE(s) learns to recover the discrete state with maximum a posteriori probability from the noisy state. Simulations show that both implementations of DOLCE lead to a significant improvement in generalization performance over earlier neural net approaches to FSM induction. The idea of adaptive quantization is not just applicable to DOLCE but can be applied to other domains as well.

Journal Article↗