Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Detection, identification, and estimation of biological aerosols and vapors with a Fourier-transform infrared spectrometer.

Two experiments were conducted with a Fourier-transform infrared (FTIR) spectrometer. The purpose of the first experiment was to detect and identify Bacillus subtilis subsp. niger (BG) bioaerosol spores and kaolin dust in an open-air release for which the thermal contrast between the aerosol temperature and background brightness temperature is small. The second experiment estimated the concentration of a small amount of triethyl phosphate (TEP) vapor in a closed chamber in which an external blackbody radiation source was used and where the thermal contrast was large. The deduced BG (TEP) extinction spectrum (identification) showed an excellent match to the library BG (TEP) extinction spectrum. Analysis of the time sequence of the measurements coincided well with the presence (detection) of the BG during the measurements, and the estimated concentration of time-dependent TEP vapor was excellent. The data were analyzed with hyperspectral detection, identification, and estimation algorithms. The algorithms were based on radiative transfer theory and statistical signal-processing methods. A subspace orthogonal projection operator was used to statistically subtract the large thermal background contribution to the measurements, and a robust maximum-likelihood solution was used to deduce the target (aerosol or vapor cloud) spectrum and estimate its mass-column concentration. A Gaussian-mixture probability model for the deduced mass-column concentration was computed with an expectation-maximization algorithm to produce the detection threshold, the probability of detection, and the probability of false alarm. The results of this study are encouraging, as they suggest for the first time to the authors' knowledge the feasibility of detecting biological aerosols with passive FTIR sensors.

Journal Article↗

A constrained EM algorithm for independent component analysis.

We introduce a novel way of performing independent component analysis using a constrained version of the expectation-maximization (EM) algorithm. The source distributions are modeled as D one-dimensional mixtures of gaussians. The observed data are modeled as linear mixtures of the sources with additive, isotropic noise. This generative model is fit to the data using constrained EM. The simpler "soft-switching" approach is introduced, which uses only one parameter to decide on the sub- or supergaussian nature of the sources. We explain how our approach relates to independent factor analysis.

Algorithms↗

Vascular segmentation of phase contrast magnetic resonance angiograms based on statistical mixture modeling and local phase coherence.

In this paper, we present an approach to segmenting the brain vasculature in phase contrast magnetic resonance angiography (PC-MRA). According to our prior work, we can describe the overall probability density function of a PC-MRA speed image as either a Maxwell-uniform (MU) or Maxwell-Gaussian-uniform (MGU) mixture model. An automatic mechanism based on Kullback-Leibler divergence is proposed for selecting between the MGU and MU models given a speed image volume. A coherence measure, namely local phase coherence (LPC), which incorporates information about the spatial relationships between neighboring flow vectors, is defined and shown to be more robust to noise than previously described coherence measures. A statistical measure from the speed images and the LPC measure from the phase images are combined in a probabilistic framework, based on the maximum a posteriori method and Markov random fields, to estimate the posterior probabilities of vessel and background for classification. It is shown that segmentation based on both measures gives a more accurate segmentation than using either speed or flow coherence information alone. The proposed method is tested on synthetic, flow phantom and clinical datasets. The results show that the method can segment normal vessels and vascular regions with relatively low flow rate and low signal-to-noise ratio, e.g., aneurysms and veins.

Algorithms↗

Estimation of linear mixed models with a mixture of distribution for the random effects.

The aim of this paper is to propose an algorithm to estimate linear mixed model when random effect distribution is a mixture of Gaussians. This heterogeneous linear mixed model relaxes the classical Gaussian assumption for the random effects and, when used for longitudinal data, can highlight distinct patterns of evolution. The observed likelihood is maximized using a Marquardt algorithm instead of the EM algorithm which is frequently used for mixture models. Indeed, the EM algorithm is computationally expensive and does not provide good convergence criteria nor direct estimates of the variance of the parameters. The proposed method also allows to classify subjects according to the estimated profiles by computing posterior probabilities of belonging to each component. The use of heterogeneous linear mixed model is illustrated through a study of the different patterns of cognitive evolution in the elderly. HETMIXLIN is a free Fortran90 program available on the web site: http://www.isped.u-bordeaux2.fr.

Aged↗

Statistical approach to segmentation of single-channel cerebral MR images.

A statistical model is presented that represents the distributions of major tissue classes in single-channel magnetic resonance (MR) cerebral images. Using the model, cerebral images are segmented into gray matter, white matter, and cerebrospinal fluid (CSF). The model accounts for random noise, magnetic field inhomogeneities, and biological variations of the tissues. Intensity measurements are modeled by a finite Gaussian mixture. Smoothness and piecewise contiguous nature of the tissue regions are modeled by a three-dimensional (3-D) Markov random field (MRF). A segmentation algorithm, based on the statistical model, approximately finds the maximum a posteriori (MAP) estimation of the segmentation and estimates the model parameters from the image data. The proposed scheme for segmentation is based on the iterative conditional modes (ICM) algorithm in which measurement model parameters are estimated using local information at each site, and the prior model parameters are estimated using the segmentation after each cycle of iterations. Application of the algorithm to a sample of clinical MR brain scans, comparisons of the algorithm with other statistical methods, and a validation study with a phantom are presented. The algorithm constitutes a significant step toward a complete data driven unsupervised approach to segmentation of MR images in the presence of the random noise and intensity inhomogeneities.

Adolescent↗

Modeling birthweight and gestational age distributions: additive vs. multiplicative processes.

Researchers have traditionally employed Gaussian distributions to model quantitative biological traits. Recently, mixtures of Gaussian distributions have begun to be used as well. However, there are many alternatives to the Gaussian distribution. From a theoretical perspective, the lognormal distribution is as applicable as the Gaussian (both are justified on the basis of the Central Limit Theorem). Here, the utility of mixtures of Gaussians and lognormals for describing birthweight and gestational age distributions are compared. This is carried out within the context of the hybrid-lognormal distribution, in which the Gaussian and lognormal are special cases. The data consists of African American births (1985-1988) and European American births (1988) in the state of New York. The results suggest that of the conventional distributions, a mixture of two Gaussians generally provides the best fit to birthweight and gestational age. However, in the case of birthweight a two-component hybrid-lognormal fits better than any of the simpler models. This may be due to a feature of the hybrid-lognormal distribution that can be interpreted as maternal constraints on fetal development.

Birth Weight↗

Gaussian mixture clustering and imputation of microarray data.

MOTIVATION: In microarray experiments, missing entries arise from blemishes on the chips. In large-scale studies, virtually every chip contains some missing entries and more than 90% of the genes are affected. Many analysis methods require a full set of data. Either those genes with missing entries are excluded, or the missing entries are filled with estimates prior to the analyses. This study compares methods of missing value estimation. RESULTS: Two evaluation metrics of imputation accuracy are employed. First, the root mean squared error measures the difference between the true values and the imputed values. Second, the number of mis-clustered genes measures the difference between clustering with true values and that with imputed values; it examines the bias introduced by imputation to clustering. The Gaussian mixture clustering with model averaging imputation is superior to all other imputation methods, according to both evaluation metrics, on both time-series (correlated) and non-time series (uncorrelated) data sets.

Algorithms↗

Stability of thin polymer films: influence of solvents.

The interface and surface properties and the wetting behavior of polymer-solvent mixtures are investigated using Monte Carlo simulations and self-consistent field calculations. We carry out Monte Carlo simulations in the framework of a coarse-grained bead-spring model using short chains (oligomers) of N(P)=5 beads and a monomeric solvent, N(S)=1. The self-consistent field calculations are based on a simple phenomenological equation of state for compressible binary mixtures and we employ Gaussian chain model. The bulk behavior of the polymer-solvent mixture belongs to type III in the classification of van Konynenburg and Scott [Phil. Trans. R. Soc. London, Ser. A 298, 495 (1980)]. It is characterized by a triple line on which the polymer-liquid coexists with solvent-vapor and a solvent-rich liquid. The solvent is not homogeneously distributed across the dense polymer film but tends to accumulate at the surface and the polymer-vapor interface. This solvent enrichment at the interface and surface becomes more pronounced upon increasing the vapor pressure and alters the surface and interface tensions. This effect gives rise to a nonmonotonic dependence of the contact angle on the vapor pressure and one might observe reentrant wetting. The results of the Monte Carlo simulations and the self-consistent field calculations qualitatively agree. The profiles of drops are investigated by Monte Carlo simulations and a pronounced solvent enrichment is observed at the wedge formed by the substrate and the liquid-vapor interface at the three-phase contact line.

Journal Article↗

Genetic analysis of somatic cell scores in US Holsteins with a Bayesian mixture model.

The objective of this study was to apply finite mixture models to field data for somatic cell scores (SCS) for estimation of genetic parameters. Data were approximately 170,000 test-day records for SCS from first-parity Holstein cows in Wisconsin. Five different models of increasing level of complexity were fitted. Model 1 was the standard single-component model, and the others were 2-component Gaussian mixtures consisting of similar but distinct linear models. All mixture models (i.e., 2 to 5) included separate means for the 2 components. Model 2 assumed entirely homogeneous variances for both components. Models 3 and 4 assumed heterogeneous variances for either residual (model 3) or genetic and permanent environmental variances (model 4). Model 5 was the most complex, in which variances of all random effects were allowed to vary across components. A Bayesian approach was applied and Gibbs sampling was used to obtain posterior estimates. Five chains of 205,000 cycles were generated for each model. Estimates of variance components were based on posterior means. Models were compared by use of the deviance information criterion. Based on the deviance information criterion, all mixture models were superior to the linear model for analysis of SCS. The best model was one in which genetic and PE variances were heterogeneous, but residual variances were homogeneous. The genetic analysis suggested that SCS in healthy and infected cattle are different traits, because the genetic correlation between SCS in the 2 components of 0.13 was significantly different from unity.

Animals↗

Independent factor analysis.

We introduce the independent factor analysis (IFA) method for recovering independent hidden sources from their observed mixtures. IFA generalizes and unifies ordinary factor analysis (FA), principal component analysis (PCA), and independent component analysis (ICA), and can handle not only square noiseless mixing but also the general case where the number of mixtures differs from the number of sources and the data are noisy. IFA is a two-step procedure. In the first step, the source densities, mixing matrix, and noise covariance are estimated from the observed data by maximum likelihood. For this purpose we present an expectation-maximization (EM) algorithm, which performs unsupervised learning of an associated probabilistic model of the mixing situation. Each source in our model is described by a mixture of gaussians; thus, all the probabilistic calculations can be performed analytically. In the second step, the sources are reconstructed from the observed data by an optimal nonlinear estimator. A variational approximation of this algorithm is derived for cases with a large number of sources, where the exact algorithm becomes intractable. Our IFA algorithm reduces to the one for ordinary FA when the sources become gaussian, and to an EM algorithm for PCA in the zero-noise limit. We derive an additional EM algorithm specifically for noiseless IFA. This algorithm is shown to be superior to ICA since it can learn arbitrary source densities from the data. Beyond blind separation, IFA can be used for modeling multidimensional data by a highly constrained mixture of gaussians and as a tool for nonlinear signal encoding.

Algorithms↗

Mixed Bayesian networks: a mixture of Gaussian distributions.

Mixed Bayesian networks are probabilistic models associated with a graphical representation, where the graph is directed and the random variables are discrete or continuous. We propose a comprehensive method for estimating the density functions of continuous variables, using a graph structure and a set of samples. The principle of the method is to learn the shape of densities from a sample of continuous variables. The densities are approximated by a mixture of Gaussian distributions. The estimation algorithm is a stochastic version of the Expectation Maximization algorithm (Stochastic EM algorithm). The inference algorithm corresponding to our model is a variant of junction three method, adapted to our specific case. The approach is illustrated by a simulated example from the domain of pharmacokinetics. Tests show that the true distributions seem sufficiently fitted for practical application.

Algorithms↗

Retinal afferents to the dorsal raphe nucleus in rats and Mongolian gerbils.

A direct pathway from the retina to the dorsal raphe nucleus (DRN) has been demonstrated in both albino rats and Mongolian gerbils. Following intraocular injection of cholera toxin subunit B (CTB), a diffuse stream of CTB-positive, fine-caliber optic axons emerged from the optic tract at the level of the pretectum/anterior mesencephalon. In gerbils, CTB-positive axons descended ventromedially into the periaqueductal gray, moving caudally and arborizing extensively throughout the DRN. In rats, the retinal-DRN projection comprised fewer, but larger caliber, axons, which arborized in a relatively restricted region of the lateral and ventral DRN. Following injection of CTB into the lateral DRN, retrogradely labeled ganglion cells (GCs) were observed in whole-mount retinas of both species. In gerbils, CTB-positive GCs were distributed over the entire retina, and a nearest-neighbor analysis of CTB-positive GCs showed significant regularity (nonrandomness) in their distribution. The overall distribution of gerbil GC soma diameters ranged from 8 to 22 micrometer and was skewed slightly towards the larger soma diameters. Based on an adaptive mixtures model statistical analysis, two Gaussian distributions appeared to comprise the total GC distribution, with mean soma diameters of 13 (SEM +/-1.7) micrometer, and 17 (SEM +/-1.5) micrometer, respectively. In rats, many fewer CTB-positive GCs were labeled following CTB injections into the lateral DRN, and nearly all occurred in the inferior retina. The total distribution of rat GC soma diameters was similar to that in gerbils and also was skewed towards the larger soma diameters. Major differences observed in the extent and configuration of the retinal-DRN pathway may be related to the diurnal/crepuscular vs. nocturnal habits of these two species.

Animals↗

A robust method for extraction and automatic segmentation of brain images.

A new protocol is introduced for brain extraction and automatic tissue segmentation of MR images. For the brain extraction algorithm, proton density and T2-weighted images are used to generate a brain mask encompassing the full intracranial cavity. Segmentation of brain tissues into gray matter (GM), white matter (WM), and cerebral spinal fluid (CSF) is accomplished on a T1-weighted image after applying the brain mask. The fully automatic segmentation algorithm is histogram-based and uses the Expectation Maximization algorithm to model a four-Gaussian mixture for both global and local histograms. The means of the local Gaussians for GM, WM, and CSF are used to set local thresholds for tissue classification. Reproducibility of the extraction procedure was excellent, with average variation in intracranial capacity (TIC) of 0.13 and 0.66% TIC in 12 healthy normal and 33 Alzheimer brains, respectively. Repeatability of the segmentation algorithm, tested on healthy normal images, indicated scan-rescan differences in global tissue volumes of less than 0.30% TIC. Reproducibility at the regional level was established by comparing segmentation results within the 12 major Talairach subdivisions. Accuracy of the algorithm was tested on a digital brain phantom, and errors were less than 1% of the phantom volume. Maximal Type I and Type II classification errors were low, ranging between 2.2 and 4.3% of phantom volume. The algorithm was also insensitive to variation in parameter initialization values. The protocol is robust, fast, and its success in segmenting normal as well as diseased brains makes it an attractive clinical application.

Adult↗

Comparative assessment of statistical brain MR image segmentation algorithms and their impact on partial volume correction in PET.

Magnetic resonance imaging (MRI)-guided partial volume effect correction (PVC) in brain positron emission tomography (PET) is now a well-established approach to compensate the large bias in the estimate of regional radioactivity concentration, especially for small structures. The accuracy of the algorithms developed so far is, however, largely dependent on the performance of segmentation methods partitioning MRI brain data into its main classes, namely gray matter (GM), white matter (WM), and cerebrospinal fluid (CSF). A comparative evaluation of three brain MRI segmentation algorithms using simulated and clinical brain MR data was performed, and subsequently their impact on PVC in 18F-FDG and 18F-DOPA brain PET imaging was assessed. Two algorithms, the first is bundled in the Statistical Parametric Mapping (SPM2) package while the other is the Expectation Maximization Segmentation (EMS) algorithm, incorporate a priori probability images derived from MR images of a large number of subjects. The third, here referred to as the HBSA algorithm, is a histogram-based segmentation algorithm incorporating an Expectation Maximization approach to model a four-Gaussian mixture for both global and local histograms. Simulated under different combinations of noise and intensity non-uniformity, MR brain phantoms with known true volumes for the different brain classes were generated. The algorithms' performance was checked by calculating the kappa index assessing similarities with the "ground truth" as well as multiclass type I and type II errors including misclassification rates. The impact of image segmentation algorithms on PVC was then quantified using clinical data. The segmented tissues of patients' brain MRI were given as input to the region of interest (RoI)-based geometric transfer matrix (GTM) PVC algorithm, and quantitative comparisons were made. The results of digital MRI phantom studies suggest that the use of HBSA produces the best performance for WM classification. For GM classification, it is suggested to use the EMS. Segmentation performed on clinical MRI data show quite substantial differences, especially when lesions are present. For the particular case of PVC, SPM2 and EMS algorithms show very similar results and may be used interchangeably. The use of HBSA is not recommended for PVC. The partial volume corrected activities in some regions of the brain show quite large relative differences when performing paired analysis on 2 algorithms, implying a careful choice of the segmentation algorithm for GTM-based PVC.

Algorithms↗

Analysis of vocal tract characteristics for near-term suicidal risk assessment.

OBJECTIVES: Among the many clinical decisions that psychiatrists must make, assessment of a patient's risk of committing suicide is definitely among the most important, complex and demanding. One of the authors reviewing his clinical experience observed that successful predictions of suicidality were often based on the patient's voice independent of content. The voices of suicidal patients exhibited unique qualities, which distinguished them from non-suicidal patients. In this study we investigated the discriminating power of lower order mel-cepstral coefficients among suicidal, major depressed, and non-suicidal patients. METHODS: Our sample consisted of 10 near-term suicidal patients, 10 major depressed patients, and 10 non-depressed control subjects. Gaussian mixtures were employed to model the class distributions of the extracted features. RESULTS AND CONCLUSIONS: As a result of two-sample ML classification analyses, first four mel-cepstral coefficients yielded exceptional classification performance with correct classification scores of 80% between near-term suicidal patients and non-depressed controls, 75% between depressed patients and non-depressed controls, and 80% between near-term suicidal patients and depressed patients.

Biomedical Engineering↗

Singularities affect dynamics of learning in neuromanifolds.

The parameter spaces of hierarchical systems such as multilayer perceptrons include singularities due to the symmetry and degeneration of hidden units. A parameter space forms a geometrical manifold, called the neuromanifold in the case of neural networks. Such a model is identified with a statistical model, and a Riemannian metric is given by the Fisher information matrix. However, the matrix degenerates at singularities. Such a singular structure is ubiquitous not only in multilayer perceptrons but also in the gaussian mixture probability densities, ARMA time-series model, and many other cases. The standard statistical paradigm of the Cramér-Rao theorem does not hold, and the singularity gives rise to strange behaviors in parameter estimation, hypothesis testing, Bayesian inference, model selection, and in particular, the dynamics of learning from examples. Prevailing theories so far have not paid much attention to the problem caused by singularity, relying only on ordinary statistical theories developed for regular (nonsingular) models. Only recently have researchers remarked on the effects of singularity, and theories are now being developed. This article gives an overview of the phenomena caused by the singularities of statistical manifolds related to multilayer perceptrons and gaussian mixtures. We demonstrate our recent results on these problems. Simple toy models are also used to show explicit solutions. We explain that the maximum likelihood estimator is no longer subject to the gaussian distribution even asymptotically, because the Fisher information matrix degenerates, that the model selection criteria such as AIC, BIC, and MDL fail to hold in these models, that a smooth Bayesian prior becomes singular in such models, and that the trajectories of dynamics of learning are strongly affected by the singularity, causing plateaus or slow manifolds in the parameter space. The natural gradient method is shown to perform well because it takes the singular geometrical structure into account. The generalization error and the training error are studied in some examples.

Journal Article↗

Computerized radiographic mass detection--part I: Lesion site selection by morphological enhancement and contextual segmentation.

This paper presents a statistical model supported approach for enhanced segmentation and extraction of suspicious mass areas from mammographic images. With an appropriate statistical description of various discriminate characteristics of both true and false candidates from the localized areas, an improved mass detection may be achieved in computer-assisted diagnosis (CAD). In this study, one type of morphological operation is derived to enhance disease patterns of suspected masses by cleaning up unrelated background clutters, and a model-based image segmentation is performed to localize the suspected mass areas using stochastic relaxation labeling scheme. We discuss the importance of model selection when a finite generalized Gaussian mixture is employed, and use the information theoretic criteria to determine the optimal model structure and parameters. Examples are presented to show the effectiveness of the proposed methods on mass lesion enhancement and segmentation when applied to mammographical images. Experimental results demonstrate that the proposed method achieves a very satisfactory performance as a preprocessing procedure for mass detection in CAD.

Breast Neoplasms↗

Style consistent classification of isogenous patterns.

In many applications of pattern recognition, patterns appear together in groups (fields) that have a common origin. For example, a printed word is usually a field of character patterns printed in the same font. A common origin induces consistency of style in features measured on patterns. The features of patterns co-occurring in a field are statistically dependent because they share the same, albeit unknown, style. Style constrained classifiers achieve higher classification accuracy by modeling such dependence among patterns in a field. Effects of style consistency on the distributions of field-features (concatenation of pattern features) can be modeled by hierarchical mixtures. Each field derives from a mixture of styles, while, within a field, a pattern derives from a class-style conditional mixture of Gaussians. Based on this model, an optimal style constrained classifier processes entire fields of patterns rendered in a consistent but unknown style. In a laboratory experiment, style constrained classification reduced errors on fields of printed digits by nearly 25 percent over singlet classifiers. Longer fields favor our classification method because they furnish more information about the underlying style.

Algorithms↗