Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bayesian feature and model selection for Gaussian mixture models.

We present a Bayesian method for mixture model training that simultaneously treats the feature selection and the model selection problem. The method is based on the integration of a mixture model formulation that takes into account the saliency of the features and a Bayesian approach to mixture learning that can be used to estimate the number of mixture components. The proposed learning algorithm follows the variational framework and can simultaneously optimize over the number of components, the saliency of the features, and the parameters of the mixture model. Experimental results using high-dimensional artificial and real data illustrate the effectiveness of the method.

Algorithms↗

Gaussian mixture modeling of alpha-helix subclasses: structure and sequence variations.

Classification of helical structures and identification of class specific sequence features is of interest for protein structure modeling. We use geometric invariant based method to first select helix-like local conformations. These conformations are mapped in a principal component space and subjected to Gaussian mixture modeling. The largest Gaussian corresponds to the regular alpha-helix. Kinked helix and curved helix appear as a separate gaussians. Class conditional, position specific amino acid propensity analysis reveals striking difference among the three classes. In regular helix, proline propensity is significant only in the beginning and low in the rest of the region regardless of length of the helix. In kinked helix, the proline propensity has a sharp peak at the helix center, while in the curved helix, the proline propensity has a broad peak in the middle region.

Amino Acids↗

Variational learning for Gaussian mixture models.

This paper proposes a joint maximum likelihood and Bayesian methodology for estimating Gaussian mixture models. In Bayesian inference, the distributions of parameters are modeled, characterized by hyperparameters. In the case of Gaussian mixtures, the distributions of parameters are considered as Gaussian for the mean, Wishart for the covariance, and Dirichlet for the mixing probability. The learning task consists of estimating the hyperparameters characterizing these distributions. The integration in the parameter space is decoupled using an unsupervised variational methodology entitled variational expectation-maximization (VEM). This paper introduces a hyperparameter initialization procedure for the training algorithm. In the first stage, distributions of parameters resulting from successive runs of the expectation-maximization algorithm are formed. Afterward, maximum-likelihood estimators are applied to find appropriate initial values for the hyperparameters. The proposed initialization provides faster convergence, more accurate hyperparameter estimates, and better generalization for the VEM training algorithm. The proposed methodology is applied in blind signal detection and in color image segmentation.

Algorithms↗

Multiscale fragile watermarking based on the Gaussian mixture model.

In this paper, a new multiscale fragile watermarking scheme based on the Gaussian mixture model (GMM) is presented. First, a GMM is developed to describe the statistical characteristics of images in the wavelet domain and an expectation-maximization algorithm is employed to identify GMM model parameters. With wavelet multiscale subspaces being divided into watermarking blocks, the GMM model parameters of different watermarking blocks are adjusted to form certain relationships, which are employed for the presented new fragile watermarking scheme for authentication. An optimal watermark embedding method is developed to achieve minimum watermarking distortion. A secret embedding key is designed to securely embed the fragile watermarks so that the new method is robust to counterfeiting, even when the malicious attackers are fully aware of the watermark embedding algorithm. It is shown that the presented new method can securely embed a message bit stream, such as personal signatures or copyright logos, into a host image as fragile watermarks. Compared with conventional fragile watermark techniques, this new statistical model based method modifies only a small amount of image data such that the distortion on the host image is imperceptible. Meanwhile, with the embedded message bits spreading over the entire image area through the statistical model, the new method can detect and localize image tampering. Besides, the new multiscale implementation of fragile watermarks based on the presented method can help distinguish some normal image operations such as JPEG compression from malicious image attacks and, thus, can be used for semi-fragile watermarking.

Algorithms↗

Automated defect detection system using wavelet packet frame and Gaussian mixture model.

This paper proposes an approach for automated defect detection in homogeneous textiles using texture analysis. The texture features are extracted by the wavelet packet frame decomposition followed by the Karhunen-Loève transform. The texture feature vector for each pixel is used as an input to a Gaussian mixture model that determines whether or not each pixel is defective. The parameters of the Gaussian mixture model are estimated with nondefective textile images in supervised defect detection. An approach for unsupervised defect detection is also presented that can identify the heterogeneous subblocks on the basis of the Kullback-Leibler divergence between two Gaussian mixtures. The proposed method was evaluated on 25 different homogeneous textile image pairs, one of each pair with a defect and the other with no defect, and was compared with existing methods using texture analysis. The experimental results yielded visually good segmentation and an excellent detection rate with a low false alarm rate for both supervised and unsupervised defect detection. This confirms the validity of the proposed approach for automated defect detection and localization.

Journal Article↗

Constrained Gaussian mixture model framework for automatic segmentation of MR brain images.

An automated algorithm for tissue segmentation of noisy, low-contrast magnetic resonance (MR) images of the brain is presented. A mixture model composed of a large number of Gaussians is used to represent the brain image. Each tissue is represented by a large number of Gaussian components to capture the complex tissue spatial layout. The intensity of a tissue is considered a global feature and is incorporated into the model through tying of all the related Gaussian parameters. The expectation-maximization (EM) algorithm is utilized to learn the parameter-tied, constrained Gaussian mixture model. An elaborate initialization scheme is suggested to link the set of Gaussians per tissue type, such that each Gaussian in the set has similar intensity characteristics with minimal overlapping spatial supports. Segmentation of the brain image is achieved by the affiliation of each voxel to the component of the model that maximized the a posteriori probability. The presented algorithm is used to segment three-dimensional, T1-weighted, simulated and real MR images of the brain into three different tissues, under varying noise conditions. Results are compared with state-of-the-art algorithms in the literature. The algorithm does not use an atlas for initialization or parameter learning. Registration processes are therefore not required and the applicability of the framework can be extended to diseased brains and neonatal brains.

Algorithms↗

Genetic-based EM algorithm for learning Gaussian mixture models.

We propose a genetic-based expectation-maximization (GA-EM) algorithm for learning Gaussian mixture models from multivariate data. This algorithm is capable of selecting the number of components of the model using the minimum description length (MDL) criterion. Our approach benefits from the properties of Genetic algorithms (GA) and the EM algorithm by combination of both into a single procedure. The population-based stochastic search of the GA explores the search space more thoroughly than the EM method. Therefore, our algorithm enables escaping from local optimal solutions since the algorithm becomes less sensitive to its initialization. The GA-EM algorithm is elitist which maintains the monotonic convergence property of the EM algorithm. The experiments on simulated and real data show that the GA-EM outperforms the EM method since: 1) We have obtained a better MDL score while using exactly the same termination condition for both algorithms. 2) Our approach identifies the number of components which were used to generate the underlying data more often than the EM algorithm.

Algorithms↗

Identification, segmentation, and image property study of acute infarcts in diffusion-weighted images by using a probabilistic neural network and adaptive Gaussian mixture model.

RATIONALE AND OBJECTIVES: Accurate identification of infarcted regions of the brain is critical in management of stroke patients. An efficient and fast method for identification and segmentation of infarcts in the diffusion-weighted images (DWI) is proposed. MATERIALS AND METHODS: Thirteen stroke patients were studied. DWI scans were acquired with a slice thickness of 5 mm. We have used a probabilistic neural network for selecting infarct slices and an adaptive (two-level) Gaussian mixture model for segmentation of the infarcts. Statistical analysis, such as identification of distribution, first-order statistics calculation, and receiver operating characteristic curve analysis, was performed. RESULTS: The average dice index is about 0.6, and average sensitivity and specificity are about 81% and 99%, respectively. The value of sensitivity and dice index are influenced by the number of false positives and false negatives. Because artifacts and infarcts have similar imaging characteristics, it is difficult to completely eliminate the artifacts. The accuracy of localization is nearly 100% as there were only two false-positive and three false-negative slices of all 381 slices. The algorithm takes about 1 minute in the Matlab computing environment to process a volume. CONCLUSION: A method to localize and segment the acute brain infarcts is proposed. The method aids the clinician in reducing the time needed to localize and segment the infarcts. The speed of localization and segmentation can be enhanced further by implementing the algorithm in VC++ and using fast algorithms for selection of Gaussian mixture model parameters.

Algorithms↗

A Gaussian mixture model based classification scheme for myoelectric control of powered upper limb prostheses.

This paper introduces and evaluates the use of Gaussian mixture models (GMMs) for multiple limb motion classification using continuous myoelectric signals. The focus of this work is to optimize the configuration of this classification scheme. To that end, a complete experimental evaluation of this system is conducted on a 12 subject database. The experiments examine the GMMs algorithmic issues including the model order selection and variance limiting, the segmentation of the data, and various feature sets including time-domain features and autoregressive features. The benefits of postprocessing the results using a majority vote rule are demonstrated. The performance of the GMM is compared to three commonly used classifiers: a linear discriminant analysis, a linear perceptron network, and a multilayer perceptron neural network. The GMM-based limb motion classification system demonstrates exceptional classification accuracy and results in a robust method of motion classification with low computational load.

Algorithms↗

Gaussian mixture models of ECoG signal features for improved detection of epileptic seizures.

PURPOSE: To investigate the potential for improving the performance of the Osorio-Frei seizure detection algorithm (OFA) by incorporating multiple FIR filters operating in parallel and Gaussian mixture models (GMM) for ECoG features distributions, thus creating "hybrid" system. METHODS: The "hybrid" algorithm decomposes the signal into four subbands, using wavelets, after which relevant features are extracted for each subband. Following these steps, multivariate GMM are developed for seizure and non-seizure states, using training segments. State classification is based on thresholding of the likelihood ratio of seizure vs. non-seizure data. Multiple comparisons are performed between this "hybrid" and a modified version of the OFA suitable for this purpose, using as indices false positives (FP), false negatives (FN) and speed of detection. RESULTS: GMM improved speed of detection over the modified OFA at negligible FP levels. The average detection delay from expert visually placed electrographic onset over all seizures was reduced from 4.8 s for modified OFA to 1.8 s for GMM (p < 0.002) Individualized training by subject proved superior to group-based training. CONCLUSIONS: This work introduces multi-feature extraction from ECoG signals together with use of Gaussian mixtures to model them, as tools to improve automated seizure detection. At the clinical level, this approach appears to increase warning time and with it the window during which safety measures and seizure blockage may be implemented, at an affordable computational cost and with negligible FP rate.

Algorithms↗

Efficient greedy learning of gaussian mixture models.

This article concerns the greedy learning of gaussian mixtures. In the greedy approach, mixture components are inserted into the mixture one after the other. We propose a heuristic for searching for the optimal component to insert. In a randomized manner, a set of candidate new components is generated. For each of these candidates, we find the locally optimal new component and insert it into the existing mixture. The resulting algorithm resolves the sensitivity to initialization of state-of-the-art methods, like expectation maximization, and has running time linear in the number of data points and quadratic in the (final) number of mixture components. Due to its greedy nature, the algorithm can be particularly useful when the optimal number of mixture components is unknown. Experimental results comparing the proposed algorithm to other methods on density estimation and texture segmentation are provided.

Algorithms↗

Clustering protein sequence and structure space with infinite Gaussian mixture models.

We describe a novel approach to the problem of automatically clustering protein sequences and discovering protein families, subfamilies etc., based on the theory of infinite Gaussian mixtures models. This method allows the data itself to dictate how many mixture components are required to model it, and provides a measure of the probability that two proteins belong to the same cluster. We illustrate our methods with application to three data sets: globin sequences, globin sequences with known three-dimensional structures and G-protein coupled receptor sequences. The consistency of the clusters indicate that our method is producing biologically meaningful results, which provide a very good indication of the underlying families and subfamilies. With the inclusion of secondary structure and residue solvent accessibility information, we obtain a classification of sequences of known structure which both reflects and extends their SCOP classifications. A supplementray web site containing larger versions of the figures is available at http://public.kgi.edu/approximately wid/PSB04/index.html

Amino Acid Sequence↗

Nonmonotonic generalization bias of Gaussian mixture models.

Theories of learning and generalization hold that the generalization bias, defined as the difference between the training error and the generalization error, increases on average with the number of adaptive parameters. This article, however, shows that this general tendency is violated for a gaussian mixture model. For temperatures just below the first symmetry breaking point, the effective number of adaptive parameters increases and the generalization bias decreases. We compute the dependence of the neural information criterion on temperature around the symmetry breaking. Our results are confirmed by numerical cross-validation experiments.

Computer Simulation↗

Birthweight by gestational age in preterm babies according to a Gaussian mixture model.

OBJECTIVE: To provide a statistically sound criterion for identifying implausibly large birthweights for gestational age. DESIGN: Review of ISTAT 1990-1994 national newborn records. SETTING: Italy POPULATION: Forty-two thousand and twenty-nine single first and second liveborn preterm babies. METHODS: Two-component Gaussian mixture models are used to describe the birthweight distributions stratified by gestational age. Implausibly large babies are identified through model-based probabilistic clustering. MAIN OUTCOME MEASURES: Gestational age misclassification and weight-for-gestational age centile curves RESULTS: Gestational age appears under-estimated by about six weeks in 12.3% of the cases. Large babies are equally present in males and females, but are more frequent in second-borns than in first-borns, even when parity-specific models are fitted. CONCLUSIONS: The approach allows for a quantification of the gestational age under-estimate error and for data correction through model-based clustering. Correct birthweight distributions and growth curves are also provided.

Birth Weight↗

Dimensionality reduction of a pathological voice quality assessment system based on Gaussian mixture models and short-term cepstral parameters.

Voice diseases have been increasing dramatically in recent times due mainly to unhealthy social habits and voice abuse. These diseases must be diagnosed and treated at an early stage, especially in the case of larynx cancer. It is widely recognized that vocal and voice diseases do not necessarily cause changes in voice quality as perceived by a listener. Acoustic analysis could be a useful tool to diagnose this type of disease. Preliminary research has shown that the detection of voice alterations can be carried out by means of Gaussian mixture models and short-term mel cepstral parameters complemented by frame energy together with first and second derivatives. This paper, using the F-Ratio and Fisher's discriminant ratio, will demonstrate that the detection of voice impairments can be performed using both mel cesptral vectors and their first derivative, ignoring the second derivative.

Computer Simulation↗

Automatic identification of Mycobacterium tuberculosis by Gaussian mixture models.

Tuberculosis and other kinds of mycobacteriosis are serious illnesses for which early diagnosis is critical for disease control. Sputum sample analysis is a common manual technique employed for bacillus detection but current sample-analysis techniques are time-consuming, very tedious, subject to poor specificity and require highly trained personnel. Image-processing and pattern-recognition techniques are appropriate tools for improving the manual screening of samples. Here we present a new technique for sputum image analysis that combines invariant shape features and chromatic channel thresholding. Some feature descriptors were extracted from an edited bacillus data set to characterize their shape. They were statistically represented by using a Gaussian mixture model representation and a minimal error Bayesian classification procedure was employed for the last identification stage. This technique constitutes a step towards automating the process and providing a high specificity.

Bacteriological Techniques↗

Classification of a known sequence of motions and postures from accelerometry data using adapted Gaussian mixture models.

Accelerometry shows promise in providing an inexpensive but effective means of long-term ambulatory monitoring of elderly patients. The accurate classification of everyday movements should allow such a monitoring system to exhibit greater 'intelligence', improving its ability to detect and predict falls by forming a more specific picture of the activities of a person and thereby allowing more accurate tracking of the health parameters associated with those activities. With this in mind, this study aims to develop more robust and effective methods for the classification of postures and motions from data obtained using a single, waist-mounted, triaxial accelerometer; in particular, aiming to improve the flexibility and generality of the monitoring system, making it better able to detect and identify short-duration movements and more adaptable to a specific person or device. Two movement classification methods were investigated: a rule-based Heuristic system and a Gaussian mixture model (GMM)-based system. A novel time-domain feature extraction method is proposed for the GMM system to allow better detection of short-duration movements. A method for adapting the GMMs to compensate for the problem of limited user-specific training data is also proposed and investigated. Classification performance was considered in relation to data gathered in an unsupervised, directed routine conducted in a three-month field trial involving six elderly subjects. The GMM system was found to achieve a mean accuracy of 91.3%, distinguishing between three postures (sitting, standing and lying) and five movements (sit-to-stand, stand-to-sit, lie-to-stand, stand-to-lie and walking), compared to 71.1% achieved by the Heuristic system. The adaptation method was found to offer a mean accuracy of 92.2%; a relative improvement of 20.2% over tests without subject-specific data and 4.5% over tests using only a limited amount of subject-specific data. While limited to a restricted subset of possible motions and postures, these results provide a significant step in the search for a more robust and accurate ambulatory classification system.

Acceleration↗