Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Tissue classification of noisy MR brain images using constrained GMM.

We present an automated algorithm for tissue segmentation of noisy, low contrast magnetic resonance (MR) images of the brain. We use a mixture model composed of a large number of Gaussians, with each brain tissue represented by a large number of the Gaussian components in order to capture the complex tissue spatial layout. The intensity of a tissue is considered a global feature and is incorporated into the model through parameter tying of all the related Gaussians. The EM algorithm is utilized to learn the parameter-tied Gaussian mixture model. A new initialization method is applied to guarantee the convergence of the EM algorithm to the global maximum likelihood. Segmentation of the brain image is achieved by the affiliation of each voxel to a selected tissue class. The presented algorithm is used to segment 3D, T1-weighted, simulated and real MR images of the brain into three different tissues, under varying noise conditions. Quantitative results are presented and compared with state-of-the-art results reported in the literature.

Algorithms↗

Mixture model for inferring susceptibility to mastitis in dairy cattle: a procedure for likelihood-based inference.

A Gaussian mixture model with a finite number of components and correlated random effects is described. The ultimate objective is to model somatic cell count information in dairy cattle and to develop criteria for genetic selection against mastitis, an important udder disease. Parameter estimation is by maximum likelihood or by an extension of restricted maximum likelihood. A Monte Carlo expectation-maximization algorithm is used for this purpose. The expectation step is carried out using Gibbs sampling, whereas the maximization step is deterministic. Ranking rules based on the conditional probability of membership in a putative group of uninfected animals, given the somatic cell information, are discussed. Several extensions of the model are suggested.

Algorithms↗

An application of reversible-jump Markov chain Monte Carlo to spike classification of multi-unit extracellular recordings.

Multi-electrode recordings in neural tissue contain the action potential waveforms of many closely spaced neurons. While we can observe the action potential waveforms, we cannot observe which neuron is the source for which waveform nor how many source neurons are being recorded. Current spike-sorting algorithms solve this problem by assuming a fixed number of source neurons and assigning the action potentials given this fixed number. We model the spike waveforms as an anisotropic Gaussian mixture model and present, as an alternative, a reversible-jump Markov chain Monte Carlo (MCMC) algorithm to simultaneously estimate the number of source neurons and to assign each action potential to a source. We derive this MCMC algorithm and illustrate its application using simulated three-dimensional data and real four-dimensional feature vectors extracted from tetrode recordings of rat entorhinal cortex neurons. In the analysis of the simulated data our algorithm finds the correct number of mixture components (sources) and classifies the action potential waveforms with minimal error. In the analysis of real data, our algorithm identifies clusters closely resembling those previously identified by a user-dependent graphical clustering procedure. Our findings suggest that a reversible-jump MCMC algorithm could offer a new strategy for designing automated spike-sorting algorithms.

Action Potentials↗

Deconvolution of evolutionary architecture unmasks a high-risk, subclonal-rich subtype in treatment-naive small cell lung cancer.

BACKGROUND: Intratumoral heterogeneity (ITH) drives therapeutic resistance in small cell lung cancer (SCLC). However, conventional single-sample analysis has limited horizontal, cross-patient comparisons, leaving the overarching evolutionary architecture in treatment-naive tumors poorly understood. This study aims to deconvolve these architectures to identify clinically relevant evolutionary subtypes. METHODS: We analyzed whole-exome sequencing data from 41 treatment-naive SCLC patients. To overcome the cross-patient comparability bottleneck, we developed a novel probabilistic framework using a refined Gaussian Mixture Model (GMM). This standardized subclonal structures into four hierarchical strata, enabling the identification of evolutionary subtypes via unsupervised clustering. To address the scarcity of SCLC public data, prognostic concordance was robustly explored in The Cancer Genome Atlas (TCGA) lung squamous cell carcinoma (LUSC) based on shared smoking etiology, with lung adenocarcinoma (LUAD) serving as a negative control. RESULTS: The cohort robustly segregated into "Clonal-dominant" (Group 1, n=28) and "Subclonal-rich" (Group 2, n=13) subtypes. Group 1 evolution was primarily driven by tobacco signatures (SBS4). Conversely, Group 2 exhibited late-stage acquisition of a DNA mismatch repair deficiency (MMRd) signature (SBS15), fueling trace subclonal diversification. Clinically, Group 2 demonstrated a significantly lower objective response rate (ORR) to platinum-based regimens (25.0% vs. 81.3%, P=0.02). Furthermore, the Subclonal-rich architecture independently predicted inferior overall survival (OS) [adjusted hazard ratio (adj. HR) =2.93, P=0.02], driven predominantly by limited-stage disease. Cross-cancer analysis validated this histology-dependent, high-heterogeneity adverse pattern in early-stage LUSC but not in LUAD. CONCLUSIONS: This hypothesis-generating study demonstrates that a "Subclonal-rich" architecture, driven by acquired MMRd, identifies high-risk, chemo-resistant SCLC. Our GMM approach suggests that pre-existing heterogeneity may serve as a potential, histology-dependent prognostic marker that warrants prospective validation for tailoring future therapeutic regimens.

Gaussian Mixture Model (GMM)↗

Source separation in astrophysical maps using independent factor analysis.

A microwave sky map results from a combination of signals from various astrophysical sources, such as cosmic microwave background radiation, synchrotron radiation and galactic dust radiation. To derive information about these sources, one needs to separate them from the measured maps on different frequency channels. Our insufficient knowledge of the weights to be given to the individual signals at different frequencies makes this a difficult task. Recent work on the problem led to only limited success due to ignoring the noise and to the lack of a suitable statistical model for the sources. In this paper, we derive the statistical distribution of some source realizations, and check the appropriateness of a Gaussian mixture model for them. A source separation technique, namely, independent factor analysis, has been suggested recently in the literature for Gaussian mixture sources in the presence of noise. This technique employs a three layered neural network architecture which allows a simple, hierarchical treatment of the problem. We modify the algorithm proposed in the literature to accommodate for space-varying noise and test its performance on simulated astrophysical maps. We also compare the performances of an expectation-maximization and a simulated annealing learning algorithm in estimating the mixture matrix and the source model parameters. The problem with expectation-maximization is that it does not ensure global optimization, and thus the choice of the starting point is a critical task. Indeed, we did not succeed to reach good solutions for random initializations of the algorithm. Conversely, our experiments with simulated annealing yielded initialization-independent results. The mixing matrix and the means and coefficients in the source model were estimated with a good accuracy while some of the variances of the components in the mixture model were not estimated satisfactorily.

Astronomy↗

AANN: an alternative to GMM for pattern recognition.

The objective in any pattern recognition problem is to capture the characteristics common to each class from feature vectors of the training data. While Gaussian mixture models appear to be general enough to characterize the distribution of the given data, the model is constrained by the fact that the shape of the components of the distribution is assumed to be Gaussian, and the number of mixtures are fixed a priori. In this context, we investigate the potential of non-linear models such as autoassociative neural network (AANN) models, which perform identity mapping of the input space. We show that the training error surface realized by the neural network model in the feature space is useful to study the characteristics of the distribution of the input data. We also propose a method of obtaining an error surface to match the distribution of the given data. The distribution capturing ability of AANN models is illustrated in the context of speaker verification.

Neural Networks, Computer↗

Predicting fundamental frequency from mel-frequency cepstral coefficients to enable speech reconstruction.

This work proposes a method to reconstruct an acoustic speech signal solely from a stream of mel-frequency cepstral coefficients (MFCCs) as may be encountered in a distributed speech recognition (DSR) system. Previous methods for speech reconstruction have required, in addition to the MFCC vectors, fundamental frequency and voicing components. In this work the voicing classification and fundamental frequency are predicted from the MFCC vectors themselves using two maximum a posteriori (MAP) methods. The first method enables fundamental frequency prediction by modeling the joint density of MFCCs and fundamental frequency using a single Gaussian mixture model (GMM). The second scheme uses a set of hidden Markov models (HMMs) to link together a set of state-dependent GMMs, which enables a more localized modeling of the joint density of MFCCs and fundamental frequency. Experimental results on speaker-independent male and female speech show that accurate voicing classification and fundamental frequency prediction is attained when compared to hand-corrected reference fundamental frequency measurements. The use of the predicted fundamental frequency and voicing for speech reconstruction is shown to give very similar speech quality to that obtained using the reference fundamental frequency and voicing.

Dichotic Listening Tests↗

Mixture model analysis of DNA microarray images.

In this paper, we propose a new methodology for analysis of microarray images. First, a new gridding algorithm is proposed for determining the individual spots and their borders. Then, a Gaussian mixture model (GMM) approach is presented for the analysis of the individual spot images. The main advantages of the proposed methodology are modeling flexibility and adaptability to the data, which are well-known strengths of GMM. The maximum likelihood and maximum a posteriori approaches are used to estimate the GMM parameters via the expectation maximization algorithm. The proposed approach has the ability to detect and compensate for artifacts that might occur in microarray images. This is accomplished by a model-based criterion that selects the number of the mixture components. We present numerical experiments with artificial and real data where we compare the proposed approach with previous ones and existing software tools for microarray image analysis and demonstrate its advantages.

Algorithms↗

Phonetically trained models for speaker recognition.

In this paper, a speaker recognition system that introduces acoustic information into a Gaussian mixture model (GMM)-based recognizer is presented. This is achieved by using a phonetic classifier during the training phase. The experimental results show that, while maintaining the recognition rate, the decrease in the computational load is between 65% and 80% depending on the number of mixtures of the models.

Adult↗

Unsupervised visual learning of three-dimensional objects using a modular network architecture.

This paper presents a modular network architecture that learns to cluster multiple views of multiple three-dimensional (3D) objects. The proposed network model is based on a mixture of non-linear autoencoders, which compete to encode multiple views of each 3D object. The main advantage of using a mixture of autoencoders is that it can capture multiple non-linear sub-spaces, rather than multiple centers for describing complex shapes of the view distributions. The unsupervised training algorithm is formulated within a maximum-likelihood estimation framework. The performance of the modular network model is evaluated through experiments using synthetic 3D wire-frame objects and gray-level images of real 3D objects. It is shown that the performance of the modular network model is superior to the performance of the conventional clustering algorithms, such as the K-means algorithm and the Gaussian mixture model.

Journal Article↗

Automatic scoring and quality assessment using accuracy bounds for FP-TDI SNP genotyping data.

BACKGROUND: Human diversity, namely single nucleotide polymorphisms (SNPs), is becoming a focus of biomedical research. Despite the binary nature of SNP determination, the majority of genotyping assay data need a critical evaluation for genotype calling. We applied statistical models to improve the automated analysis of 2-dimensional SNP data. METHODS: We derived several quantities in the framework of Gaussian mixture models that provide figures of merit to objectively measure the data quality. The accuracy of individual observations is scored as the probability of belonging to a certain genotype cluster, while the assay quality is measured by the overlap between the genotype clusters. RESULTS: The approach was extensively tested with a dataset of 438 nonredundant SNP assays comprising >150,000 datapoints. The performance of our automatic scoring method was compared with manual assignments. The agreement for the overall assay quality is remarkably good, and individual observations were scored differently by man and machine in 2.6% of cases, when applying stringent probability threshold values. CONCLUSION: Our definition of bounds for the accuracy for complete assays in terms of misclassification probabilities goes beyond other proposed analysis methods. We expect the scoring method to minimise human intervention and provide a more objective error estimate in genotype calling.

Algorithms↗

Using statistical image models for objective evaluation of spot detection in two-dimensional gels.

Protein spot detection is central to the analysis of two-dimensional electrophoresis gel images. There are many commercially available packages, each implementing a protein spot detection algorithm. Despite this, there have been relatively few studies comparing the performance characteristics of the different packages. This is in part due to the fact that different packages employ different sets of user-adjustable parameters. It is also partly due to the fact that the images are complex. To carry out an evaluation, "ground truth" data specifying spot position, shape and intensities needs to be defined subjectively on selected test images. We address this problem by proposing a method of evaluation using synthetic images with unambiguous interpretation. The characteristics of the spots in the synthetic images are determined from statistical models of the shape, intensity, size, spread and location of real spot data. The distribution of parameters is described using a Gaussian mixture model obtained from training images. The synthetic images allow us to investigate the effects of individual image properties, such as signal-to-noise ratios and degree of spot overlap, by measuring quantifiable outcomes, e.g. accuracy of spot position, false positive and false negative detection. We illustrate the approach by carrying out quantitative evaluations of spot detection on a number of widely used analysis packages.

Algorithms↗

Bayesian fluorescence in situ hybridisation signal classification.

Previous research has indicated the significance of accurate classification of fluorescence in situ hybridisation (FISH) signals for the detection of genetic abnormalities. Based on well-discriminating features and a trainable neural network (NN) classifier, a previous system enabled highly-accurate classification of valid signals and artefacts of two fluorophores. However, since this system employed several features that are considered independent, the naive Bayesian classifier (NBC) is suggested here as an alternative to the NN. The NBC independence assumption permits the decomposition of the high-dimensional likelihood of the model for the data into a product of one-dimensional probability densities. The naive independence assumption together with the Bayesian methodology allow the NBC to predict a posteriori probabilities of class membership using estimated class-conditional densities in a close and simple form. Since the probability densities are the only parameters of the NBC, the misclassification rate of the model is determined exclusively by the quality of density estimation. Densities are evaluated by three methods: single Gaussian estimation (SGE; parametric method), Gaussian mixture model assuming spherical covariance matrices (GMM; semi-parametric method) and kernel density estimation (KDE; non-parametric method). For low-dimensional densities, the GMM generally outperforms the KDE that tends to overfit the training set at the cost of reduced generalisation capability. But, it is the GMM that loses some accuracy when modelling higher-dimensional densities due to the violation of the assumption of spherical covariance matrices when dependent features are added to the set. Compared with these two methods, the SGE and NN provide inferior and superior performance, respectively. However, the NBC avoids the intensive training and optimisation required for the NN, demanding extensive resources and experimentation. Therefore, when supporting these two classifiers, the system enables a trade-off between the NN performance and NBC simplicity of implementation.

Algorithms↗

Variability of birth-weight distributions by sex and ethnicity: analysis using mixture models.

Birth weight is the most important proximate determinant of the level of infant mortality. However, the association between birth weight and infant mortality is not constant among populations. For example, the mortality of African American infants is lower at low birth weight but higher at high birth weight compared with European American infants. One possible explanation is that birth cohorts are heterogeneous even after controlling for birth weight, ethnicity, sex, and multiple births. The analyses presented here use Gaussian mixture models to explore the interpopulation variation in the shape of the birth-weight distribution for evidence of intrapopulation heterogeneity. The results suggest that a two-component mixture model provides an excellent description of human birth-weight distributions. Further statistical analyses of sex and ethnic differences indicate (1) that the birth-weight distributions and heterogeneity within the distribution vary between the sexes and among ethnic groups and (2) that one specific component is more closely associated with the overall level of infant mortality. The results support the hypothesis that birth cohorts can consist of two or more subpopulations at differential risk of mortality. Differences in the subpopulation composition of birth cohorts (i.e., differences in the level of heterogeneity among the various ethnic groups) might partially explain the interethnic variation in birth-weight-specific mortality. Further development of these mixture models should provide important additional information concerning the biological, environmental, and social determinants of birth weight and infant mortality.

Black or African American↗

Robust speaker's location detection in a vehicle environment using GMM models.

Abstract-Human-computer interaction (HCI) using speech communication is becoming increasingly important, especially in driving where safety is the primary concern. Knowing the speaker's location (i.e., speaker localization) not only improves the enhancement results of a corrupted signal, but also provides assistance to speaker identification. Since conventional speech localization algorithms suffer from the uncertainties of environmental complexity and noise, as well as from the microphone mismatch problem, they are frequently not robust in practice. Without a high reliability, the acceptance of speech-based HCI would never be realized. This work presents a novel speaker's location detection method and demonstrates high accuracy within a vehicle cabinet using a single linear microphone array. The proposed approach utilize Gaussian mixture models (GMM) to model the distributions of the phase differences among the microphones caused by the complex characteristic of room acoustic and microphone mismatch. The model can be applied both in near-field and far-field situations in a noisy environment. The individual Gaussian component of a GMM represents some general location-dependent but content and speaker-independent phase difference distributions. Moreover, the scheme performs well not only in nonline-of-sight cases, but also when the speakers are aligned toward the microphone array but at difference distances from it. This strong performance can be achieved by exploiting the fact that the phase difference distributions at different locations are distinguishable in the environment of a car. The experimental results also show that the proposed method outperforms the conventional multiple signal classification method (MUSIC) technique at various SNRs.

Acoustics↗

Probabilistic space-time video modeling via piecewise GMM.

In this paper, we describe a statistical video representation and modeling scheme. Video representation schemes are needed to segment a video stream into meaningful video-objects, useful for later indexing and retrieval applications. In the proposed methodology, unsupervised clustering via Gaussian mixture modeling extracts coherent space-time regions in feature space, and corresponding coherent segments (video-regions) in the video content. A key feature of the system is the analysis of video input as a single entity as opposed to a sequence of separate frames. Space and time are treated uniformly. The probabilistic space-time video representation scheme is extended to a piecewise GMM framework in which a succession of GMMs are extracted for the video sequence, instead of a single global model for the entire sequence. The piecewise GMM framework allows for the analysis of extended video sequences and the description of nonlinear, nonconvex motion patterns. The extracted space-time regions allow for the detection and recognition of video events. Results of segmenting video content into static versus dynamic video regions and video content editing are presented.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

A fragment library based on Gaussian mixtures predicting favorable molecular interactions.

Here, a protein atom-ligand fragment interaction library is described. The library is based on experimentally solved structures of protein-ligand and protein-protein complexes deposited in the Protein Data Bank (PDB) and it is able to characterize binding sites given a ligand structure suitable for a protein. A set of 30 ligand fragment types were defined to include three or more atoms in order to unambiguously define a frame of reference for interactions of ligand atoms with their receptor proteins. Interactions between ligand fragments and 24 classes of protein target atoms plus a water oxygen atom were collected and segregated according to type. The spatial distributions of individual fragment - target atom pairs were visually inspected in order to obtain rough-grained constraints on the interaction volumes. Data fulfilling these constraints were given as input to an iterative expectation-maximization algorithm that produces as output maximum likelihood estimates of the parameters of the finite Gaussian mixture models. Concepts of statistical pattern recognition and the resulting mixture model densities are used (i) to predict the detailed interactions between Chlorella virus DNA ligase and the adenine ring of its ligand and (ii) to evaluate the "error" in prediction for both the training and validation sets of protein-ligand interaction found in the PDB. These analyses demonstrate that this approach can successfully narrow down the possibilities for both the interacting protein atom type and its location relative to a ligand fragment.

Algorithms↗