Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Efficiency and robustness in the use of repeated measurements.

Efficiencies in repeated-measurements experiments are studied by the use of F statistics when appropriate and of Hotelling's T2 statistics otherwise. Bounds on the asymptotic relative efficiencies of these statistics are given and conditions are found under which one test dominates another uniformly. Similar comparisons are made between T2 and the Lawley-Hotelling T2(0) test in the analysis of repeated vector measurements. It is found that the T2 test may be considerably more efficient than its competitors, but that its efficiency cannot be substantially less. These findings are illustrated numerically by examples from the literature. It is shown further that F and T2 tests are exact for all unimodal distributions constant on ellipsoids, and that efficiency comparisons among F tests are preserved under scale mixtures of multidimensional Gaussian laws.

Humans↗

Expert opinion elicitation for assisting deep learning based Lyme disease classifier with patient data.

BACKGROUND: Diagnosing erythema migrans (EM) skin lesion, the most common early symptom of Lyme disease, using deep learning techniques can be effective to prevent long-term complications. Existing works on deep learning based EM recognition only utilizes lesion image due to the lack of a dataset of Lyme disease related images with associated patient data. Doctors rely on patient information about the background of the skin lesion to confirm their diagnosis. To assist deep learning model with a probability score calculated from patient data, this study elicited opinions from fifteen expert doctors. To the best of our knowledge, this is the first expert elicitation work to calculate Lyme disease probability from patient data. METHODS: For the elicitation process, a questionnaire with questions and possible answers related to EM was prepared. Doctors provided relative weights to different answers to the questions. We converted doctors' evaluations to probability scores using Gaussian mixture based density estimation. We exploited formal concept analysis and decision tree for elicited model validation and explanation. We also proposed an algorithm for combining independent probability estimates from multiple modalities, such as merging the EM probability score from a deep learning image classifier with the elicited score from patient data. RESULTS: We successfully elicited opinions from fifteen expert doctors to create a model for obtaining EM probability scores from patient data. CONCLUSIONS: The elicited probability score and the proposed algorithm can be utilized to make image based deep learning Lyme disease pre-scanners robust. The proposed elicitation and validation process is easy for doctors to follow and can help address related medical diagnosis problems where it is challenging to collect patient data.

Humans↗

Direct correlation functions of binary mixtures of hard Gaussian overlap molecules.

We study the direct correlation function (DCF) of a classical fluid mixture of nonspherical molecules. The components of the mixture are two types of hard ellipsoidal molecules with different elongations, interacting through the hard Gaussian overlap (HGO) model. Two different approaches are used to calculate the DCFs of this fluid, and the results are compared. Here, the Pynn approximation [J. Chem. Phys. 60, 4579 (1974)] is extended to calculate the DCF of the binary mixtures of HGO molecules, then we use a formalism based on the weighted density functional theory introduced by Chamoux and Perera [J. Chem. Phys. 104, 1493 (1996)]. These results are fairly in agreement with each other. The pressure of this system is also calculated using the Fourier zero components of the DCF. The results are in agreement with the Monte Carlo molecular simulation.

Journal Article↗

Comparison and validation of tissue modelization and statistical classification methods in T1-weighted MR brain images.

This paper presents a validation study on statistical nonsupervised brain tissue classification techniques in magnetic resonance (MR) images. Several image models assuming different hypotheses regarding the intensity distribution model, the spatial model and the number of classes are assessed. The methods are tested on simulated data for which the classification ground truth is known. Different noise and intensity nonuniformities are added to simulate real imaging conditions. No enhancement of the image quality is considered either before or during the classification process. This way, the accuracy of the methods and their robustness against image artifacts are tested. Classification is also performed on real data where a quantitative validation compares the methods' results with an estimated ground truth from manual segmentations by experts. Validity of the various classification methods in the labeling of the image as well as in the tissue volume is estimated with different local and global measures. Results demonstrate that methods relying on both intensity and spatial information are more robust to noise and field inhomogeneities. We also demonstrate that partial volume is not perfectly modeled, even though methods that account for mixture classes outperform methods that only consider pure Gaussian classes. Finally, we show that simulated data results can also be extended to real data.

Adult↗

Bayesian analysis of mixtures applied to post-synaptic potential fluctuations.

Bayesian inference techniques have been applied to the analysis of fluctuation of post-synaptic potentials in the hippocampus. The underlying statistical model assumes that the varying synaptic signals are characterized by mixtures of (unknown) numbers of individual gaussian, or normal, component distributions. Each solution consists of a group of individual components with unique mean values and relative probabilities of occurrence and a predictive probability density. The advantages of bayesian inference techniques over the alternative method of maximum likelihood estimation (MLE) of the parameters of an unknown mixture distribution include the following: (1) prior information may be incorporated in the estimation of model parameters; (2) conditional probability estimates of the number of individual components in the mixture are calculated; (3) flexibility exists in the extent to which the estimated noise standard deviation indicates the width of each component; (4) posterior distributions for component means are calculated, including measures of uncertainty about the means; and (5) probability density functions of the component distributions and the overall mixture distribution are estimated in relation to the raw grouped data, together with measures of uncertainty about these estimates. This expository report describes this novel approach to the unconstrained identification of components within a mixture, and provides demonstration of the usefulness of the technique in the context of both simulations and the analysis of distributions of synaptic potential signals.

Action Potentials↗

Continuous trees and NEVADA simulation: a quadrature approach to modeling continuous random variables in decision analysis.

This paper introduces an improved technique for modeling risk and decision problems that have continuous random variables and probabilistic dependence. Variables are modeled with mixtures of four-parameter random variables, called "continuous trees." Functions of random variables are calculated using gaussian quadrature in a manner called "Nevada simulation" (NumErical Integration of Variance And probabilistic Dependence Analyzer). This technique is compared with traditional decision-tree modeling in terms of analytic technique, solution-time complexity, and accuracy. Nevada simulation takes advantage of the probabilistic independence in a decision problem while allowing for probabilistic dependence to achieve polynomial computational-time complexity for many decision problems. It improves on the accuracy of traditional decision trees by employing larger approximations than traditional decision analysis. It improves on traditional decision analysis by modeling continuous variables with continuous, rather than discrete, distributions. A Bayesian analysis using a mixed discrete-continuous probability distribution for cigarette smoking rate is presented.

Algorithms↗

The normal distribution in scaling subjective stimulus differences: less "normal" than we think?

The Gaussian, or "normal," distribution is routinely used to model distributions of subjective quantities. This practice rests on a trust in the central limit theorem, yet this theorem does not cover the case of a mixture of Gaussian distributions with different standard deviations, which yields a distribution that is heavier tailed than the Gaussian. In psychophysical judgment experiments, this may result, for example, from fluctuating or interindividually varying attention. Two candidates for describing this commonly encountered type of distribution are the logistic distribution and the t distribution with a small number of degrees of freedom. In reanalyses of experimental data on three-category loudness comparisons, as well as in a Monte Carlo simulation, t(4) was found to model the underlying mixed inter- and intraindividual distribution of subjective loudness differences quite satisfactorily.

Humans↗

Lack of a bimodal distribution of ventricular size in schizophrenia: a Gaussian mixture analysis of 1056 cases and controls.

The finding of clinical and laboratory differences between schizophrenic patients with large and small cerebral ventricles has led to the widespread assumption that large ventricles are a marker that characterizes a subgroup of patients with schizophrenia. We reviewed all published English language ventricle-to-brain ratio (VBR) studies in which individual data points were available (schizophrenics: n = 691, medical controls; n = 205, normal volunteers: n = 160). Using a univariate normal mixture model to examine the distribution of ventricular size in each group, we found no evidence of a mixture of Gaussian distributions (i.e., "bimodality") within any of the three groups. The same analysis was then performed on the combined sample of schizophrenic patients and normal and medical controls, respectively. In each case the improvement in fit of a mixture of normal distributions compared to a single component normal distribution was significant. The data do not support the notion that ventricular enlargement is a discontinuous marker of a subtype of schizophrenia.

Analysis of Variance↗

Partial volume segmentation of brain magnetic resonance images based on maximum a posteriori probability.

Noise, partial volume (PV) effect, and image-intensity inhomogeneity render a challenging task for segmentation of brain magnetic resonance (MR) images. Most of the current MR image segmentation methods focus on only one or two of the above-mentioned effects. The objective of this paper is to propose a unified framework, based on the maximum a posteriori probability principle, by taking all these effects into account simultaneously in order to improve image segmentation performance. Instead of labeling each image voxel with a unique tissue type, the percentage of each voxel belonging to different tissues, which we call a mixture, is considered to address the PV effect. A Markov random field model is used to describe the noise effect by considering the nearby spatial information of the tissue mixture. The inhomogeneity effect is modeled as a bias field characterized by a zero mean Gaussian prior probability. The well-known fuzzy C-mean model is extended to define the likelihood function of the observed image. This framework reduces theoretically, under some assumptions, to the adaptive fuzzy C-mean (AFCM) algorithm proposed by Pham and Prince. Digital phantom and real clinical MR images were used to test the proposed framework. Improved performance over the AFCM algorithm was observed in a clinical environment where the inhomogeneity, noise level, and PV effect are commonly encountered.

Algorithms↗

On-line estimation of concentration parameters in fermentation processes.

It has long been thought that bioprocess, with their inherent measurement difficulties and complex dynamics, posed almost insurmountable problems to engineers. A novel software sensor is proposed to make more effective use of those measurements that are already available, which enable improvement in fermentation process control. The proposed method is based on mixtures of Gaussian processes (GP) with expectation maximization (EM) algorithm employed for parameter estimation of mixture of models. The mixture model can alleviate computational complexity of GP and also accord with changes of operating condition in fermentation processes, i.e., it would certainly be able to examine what types of process-knowledge would be most relevant for local models' specific operating points of the process and then combine them into a global one. Demonstrated by on-line estimate of yeast concentration in fermentation industry as an example, it is shown that soft sensor based state estimation is a powerful technique for both enhancing automatic control performance of biological systems and implementing on-line monitoring and optimization.

Algorithms↗

The Rician inverse Gaussian distribution: a new model for non-Rayleigh signal amplitude statistics.

In this paper, we introduce a new statistical distribution for modeling non-Rayleigh amplitude statistics, which we have called the Rician inverse Gaussian (RiIG) distribution. It is a mixture of the Rice distribution and the inverse Gaussian distribution. The probability density function (pdf) is given in closed form as a function of three parameters. This makes the pdf very flexible in the sense that it may be fitted to a variety of shapes, ranging from the Rayleigh-shaped pdf to a noncentral chi2-shaped pdf. The theoretical basis of the new model is quite thoroughly discussed, and we also give two iterative algorithms for estimating its parameters from data. Finally, we include some modeling examples, where we have tested the ability of the distribution to represent locale amplitude histograms of linear medical ultrasound data and single-look synthetic aperture radar data. We compare the goodness of fit of the RiIG model with that of the K model, and, in most cases, the new model turns out as a better statistical model for the data. We also include a series of log-likelihood tests to evaluate the predictive performance of the proposed model.

Algorithms↗

Predicting the time to quasi-extinction for populations far below their carrying capacity.

Populations threatened by extinction are often far below their carrying capacity. A population collapse or quasi-extinction is defined to occur when the population size reaches some given lower density. If this density is chosen to be large enough for the demographic stochasticity to be ignored compared to environmental stochasticity, then the logarithm of the population size may be modelled by a Brownian motion until quasi-extinction occurs. The normal-gamma mixture of inverse Gaussian distributions can then be applied to define prediction intervals for the time to quasi-extinction in such processes. A similar mixture is used to predict the population size at a finite time for the same process provided that quasi-extinction has not occurred before that time. Stochastic simulations indicate that the coverage of the prediction interval is very close to the probability calculated theoretically. As an illustration, the method is applied to predict the time to extinction of a declining population of white stork in southwestern Germany.

Animals↗

Fetal electrocardiogram extraction by sequential source separation in the wavelet domain.

This paper addresses the problem of fetal electrocardiogram extraction using blind source separation (BSS) in the wavelet domain. A new approach is proposed, which is particularly advantageous when the mixing environment is noisy and time-varying, and that is shown, analytically and in simulation, to improve the convergence rate of the natural gradient algorithm. The distribution of the wavelet coefficients of the source signals is then modeled by a generalized Gaussian probability density, whereby in the time-scale domain the problem of selecting appropriate nonlinearities when separating mixtures of both sub- and super-Gaussian signals is mitigated, as shown by experimental results.

Algorithms↗

Discrete velocity and lattice Boltzmann models for binary mixtures of nonideal fluids.

In this paper, a discrete velocity model and a lattice Boltzmann model are proposed for binary mixtures of nonideal fluids based on the Enskog theory. The velocity space of the Enskog equation for each component is first discretized by applying a Gaussian quadrature, resulting in a discrete velocity model that can be solved by suitable numerical schemes. A lattice Boltzmann model is then derived from the discrete velocity model with a slightly modified equilibrium. The hydrodynamics of each model are also derived through the Chapmann-Enskog procedure.

Journal Article↗

Improved system for object detection and star/galaxy classification via local subspace analysis.

The two traditional tasks of object detection and star/galaxy classification in astronomy can be automated by neural networks because the nature of the problems is that of pattern recognition. A typical existing system can be further improved by using one of the local Principal Component Analysis (PCA) models. Our analysis in the context of object detection and star/galaxy classification reveals that local PCA is not only superior to global PCA in feature extraction, but is also superior to gaussian mixture in clustering analysis. Unlike global PCA which performs PCA for the whole data set, local PCA applies PCA individually to each cluster of data. As a result, local PCA often outperforms global PCA for data of multi-modes. Moreover, since local PCA can effectively avoid the trouble of having to specify a large number of free elements of each covariance matrix of gaussian mixture, it can give a better description of local subspace structures of each cluster when applied on high dimensional data with small sample size. In this paper, the local PCA model proposed by Xu [IEEE Trans. Neural Networks 12 (2001) 822] under the general framework of Bayesian Ying Yang (BYY) normalization learning will be adopted. Endowed with the automatic model selection ability of BYY learning, the BYY normalization learning-based local PCA model can cope with those object detection and star/galaxy classification tasks with unknown model complexity. A detailed algorithm for implementation of the local PCA model will be proposed, and experimental results using both synthetic and real astronomical data will be demonstrated.

Astronomy↗

Brain tissue classification of magnetic resonance images using partial volume modeling.

This paper presents a fully automatic three-dimensional classification of brain tissues for Magnetic Resonance (MR) images. An MR image volume may be composed of a mixture of several tissue types due to partial volume effects. Therefore, we consider that in a brain dataset there are not only the three main types of brain tissue: gray matter, white matter, and cerebro spinal fluid, called pure classes, but also mixtures, called mixclasses. A statistical model of the mixtures is proposed and studied by means of simulations. It is shown that it can be approximated by a Gaussian function under some conditions. The D'Agostino-Pearson normality test is used to assess the risk alpha of the approximation. In order to classify a brain into three types of brain tissue and deal with the problem of partial volume effects, the proposed algorithm uses two steps: 1) segmentation of the brain into pure and mixclasses using the mixture model; 2) reclassification of the mixclasses into the pure classes using knowledge about the obtained pure classes. Both steps use Markov random field (MRF) models. The multifractal dimension, describing the topology of the brain, is added to the MRFs to improve discrimination of the mixclasses. The algorithm is evaluated using both simulated images and real MR images with different T1-weighted acquisition sequences.

Algorithms↗

Probabilistic Monte Carlo based mapping of cerebral connections utilising whole-brain crossing fibre information.

A methodology is presented for estimation of a probability density function of cerebral fibre orientations when one or two fibres are present in a voxel. All data are acquired on a clinical MR scanner, using widely available acquisition techniques. The method models measurements of water diffusion in a single fibre by a Gaussian density function and in multiple fibres by a mixture of Gaussian densities. The effects of noise on complex MR diffusion weighted data are explicitly simluated and parameterised. This information is used for standard and Monte Carlo streamline methods. Deterministic and probabilistic maps of anatomical voxel scale connectivity between brain regions are generated.

Algorithms↗

Estimating heterogeneity in random effects models for longitudinal data.

In this paper, we are interested in estimating parameters entering nonlinear mixed effects models using a likelihood maximization approach. As the accuracy of the likelihood approximation is likely to govern the quality of the derived estimates of both the distribution of the random effects and the fixed parameters, we propose a methodological approach based on the adaptive Gauss Hermite quadrature to better approximate the likelihood function. This work presents improvements of this quadrature that render it accurate and computationally efficient in the problem of likelihood approximation with, an application to mixture models, models which allow the description of coexistence of several different homogeneous subpopulations specifying the distribution of random effects as a mixture of Gaussian distributions. These improvements are based on a new choice of the scaling matrix followed by its optimisation. An application to a phase III clinical trial of an anticoagulant molecule is proposed and estimation results are compared to those obtained with the most frequently used method in population pharmacokinetic analysis. Moreover, in order to evaluate the accuracy of the estimations, an analysis of simulated pharmacokinetic data derived from the model and the a priori values of population parameters of the previous study are presented.

Algorithms↗