Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

A nonparametric approach to extract information from interspike interval data.

In this work we develop an approach to extracting information from neural spike trains. Using the expectation-maximization (EM) algorithm, interspike interval data from experiments and simulations are fitted by mixtures of distributions, including Gamma, inverse Gaussian, log-normal, and the distribution of the interspike intervals of the leaky integrate-and-fire model. In terms of the Kolmogorov-Smirnov test for goodness-of-fit, our approach is proved successful (P>0.05) in fitting benchmark data for which a classical parametric approach has been shown to fail before. In addition, we present a novel method to fit mixture models to censored data, and discuss two examples of the application of such a method, which correspond to the case of multiple-trial and multielectrode array data. A MATLAB implementation of the algorithm is available for download from .

Action Potentials↗

Prediction of peak shape as a function of retention in reversed-phase liquid chromatography.

Optimisation of the resolution of multicomponent samples in HPLC is usually carried out by changing the elution conditions and considering the variation in retention of the analytes, to which a standard peak shape is assigned. However, the change in peak shape with the composition of the mobile phase can ruin the optimisation process, yielding unexpected overlaps in the experimental chromatograms for the predicted optimum, especially for complex mixtures. The possibility of modelling peak shape, in addition to peak position, is therefore attractive. A simple modified-Gaussian model with a parabolic variance, which is a function of conventional experimental parameters: retention time (tR), peak height (H0), standard deviation at the peak maximum (sigma0), and left (A) and right (B) halfwidths, is proposed. The model is a simplification of a previous equation proposed in our laboratory. Linear and parabolic relationships were found between the peak shape parameters (sigma0), A and B) and tR, with a mean relative error of 1-5% in most cases. This error was partially due to variations in peak position and shape among injections, which in some cases were above 2%. Correlations between (sigma0, A and B) and the retention time, which is easily modelled as a function of mobile phase composition, allowed a simple and reliable prediction of chromatographic peaks. A parameter that depends on the slopes of the linear relationships for A and B versus tR is also proposed to evaluate column efficiency. The modified-Gaussian model was used to describe the peaks of six diuretics of diverse acid-base behaviour and polarity, which were eluted with 15 mobile phases where the composition was varied between 30 and 50% (v/v) acetonitrile and the pH between 3 and 7.

Chromatography, Liquid↗

Admixture analysis of age at onset in obsessive-compulsive disorder.

BACKGROUND: Age at onset (AAO) has been useful to explore the clinical, neurobiological and genetic heterogeneity of obsessive-compulsive disorder (OCD). However, none of the various thresholds of AAO used in previous studies have been validated, and it remains an unproven notion that AAO is a marker for different subtypes of OCD. If AAO is a clinical indicator of different biological subtypes, then subgroups based on distinct AAOs should have separate normal distributions as well as different clinical characteristics. METHOD: Admixture analysis was used to determine the best-fitting model for the observed AAO of 161 OCD patients. RESULTS: The observed distribution of AAO in OCD is a mixture of two Gaussian distributions with mean ages of 11.1 +/- 4.1 and 23.5 +/- 11.1 years. The first distribution, defined by early-onset OCD, had increased frequency of Tourette's syndrome and increased family history of OCD. The second distribution, defined by late-onset OCD, showed elevated prevalence of general anxiety disorder and major depressive disorder. CONCLUSIONS: These results, based on a statistically validated AAO cut-off and those of previous studies on AAO in OCD, suggest that AAO is a crucial phenotypic characteristic in understanding the genetic basis of this disorder.

Adolescent↗

On the origin of skewed distributions of spontaneous synaptic potentials in autonomic ganglia.

The histograms of spontaneous synaptic potentials at synapses in autonomic ganglia are described by distributions consisting of mixtures of Gaussians, rather than by single Gaussian distributions. The possible origin of these mixed distributions is investigated, using Monte-Carlo simulations of the action of spontaneously released units of transmitter. A single unit of acetylcholine of fixed size, released from an active zone with receptor patches both beneath and adjacent to the zone, does not give rise to the observed histograms. But if the unit is of variable size, consisting of integer multiples of smaller units, and release is from an active zone onto either the receptor patch beneath, or in addition onto adjacent patches, then the histogram is well described by a mixture of Gaussians. However, this explanation is unlikely to be correct as present evidence suggests that in most cases the released unit of transmitter saturates the postsynaptic receptor patch beneath the active zone. The final case considered is where a unit of transmitter is spontaneously released from an active zone, simultaneously with a unit in an adjacent zone less than one micron away. The histogram of potentials then conforms to those observed even when there are differences in the sizes of the receptor patches. It is suggested that this kind of release could provide an explanation for distributions of spontaneous potentials that are mixtures of Gaussians.

Animals↗

Blind source separation using temporal predictability.

A measure of temporal predictability is defined and used to separate linear mixtures of signals. Given any set of statistically independent source signals, it is conjectured here that a linear mixture of those signals has the following property: the temporal predictability of any signal mixture is less than (or equal to) that of any of its component source signals. It is shown that this property can be used to recover source signals from a set of linear mixtures of those signals by finding an un-mixing matrix that maximizes a measure of temporal predictability for each recovered signal. This matrix is obtained as the solution to a generalized eigenvalue problem; such problems have scaling characteristics of O(N3), where N is the number of signal mixtures. In contrast to independent component analysis, the temporal predictability method requires minimal assumptions regarding the probability density functions of source signals. It is demonstrated that the method can separate signal mixtures in which each mixture is a linear combination of source signals with supergaussian, subgaussian, and gaussian probability density functions and on mixtures of voices and music.

Humans↗

Joint modeling of DNA sequence and physical properties to improve eukaryotic promoter recognition.

We present an approach to integrate physical properties of DNA, such as DNA bendability or GC content, into our probabilistic promoter recognition system McPROMOTER. In the new model, a promoter is represented as a sequence of consecutive segments represented by joint likelihoods for DNA sequence and profiles of physical properties. Sequence likelihoods are modeled with interpolated Markov chains, physical properties with Gaussian distributions. The background uses two joint sequence/profile models for coding and non-coding sequences, each consisting of a mixture of a sense and an anti-sense submodel. On a large Drosophila test set, we achieved a reduction of about 30% of false positives when compared with a model solely based on sequence likelihoods.

Animals↗

Stochastic trapping in a solvable model of on-line independent component analysis.

Previous analytical studies of on-line independent component analysis (ICA) learning rules have focused on asymptotic stability and efficiency. In practice, the transient stages of learning are often more significant in determining the success of an algorithm. This is demonstrated here with an analysis of a Hebbian ICA algorithm, which can find a small number of nongaussian components given data composed of a linear mixture of independent source signals. An idealized data model is considered in which the sources comprise a number of nongaussian and gaussian sources, and a solution to the dynamics is obtained in the limit where the number of gaussian sources is infinite. Previous stability results are confirmed by expanding around optimal fixed points, where a closed-form solution to the learning dynamics is obtained. However, stochastic effects are shown to stabilize otherwise unstable suboptimal fixed points. Conditions required to destabilize one such fixed point are obtained for the case of a single nongaussian component, indicating that the initial learning rate eta required to escape successfully is very low (eta = O(N(-2)) where N is the data dimension), resulting in very slow learning typically requiring O(N(3)) iterations. Simulations confirm that this picture holds for a finite system.

Journal Article↗

SMEM algorithm for mixture models.

We present a split-and-merge expectation-maximization (SMEM) algorithm to overcome the local maxima problem in parameter estimation of finite mixture models. In the case of mixture models, local maxima often involve having too many components of a mixture model in one part of the space and too few in another, widely separated part of the space. To escape from such configurations, we repeatedly perform simultaneous split-and-merge operations using a new criterion for efficiently selecting the split-and-merge candidates. We apply the proposed algorithm to the training of gaussian mixtures and mixtures of factor analyzers using synthetic and real data and show the effectiveness of using the split-and-merge operations to improve the likelihood of both the training data and of held-out test data. We also show the practical usefulness of the proposed algorithm by applying it to image compression and pattern recognition problems.

Algorithms↗

Practical Bayesian inference using mixtures of mixtures.

Discrete mixtures of normal distributions are widely used in modeling amplitude fluctuations of electrical potentials at synapses of human and other animal nervous systems. The usual framework has independent data values yj arising as yj = mu j + xn0 + j, where the means mu j come from some discrete prior G(mu) and the unknown xno + j's and observed xj, j = 1,...,n0, are Gaussian noise terms. A practically important development of the associated statistical methods is the issue of nonnormality of the noise terms, often the norm rather than the exception in the neurological context. We have recently developed models, based on convolutions of Dirichlet process mixtures, for such problems. Explicitly, we model the noise data values xj as arising from a Dirichlet process mixture of normals, in addition to modeling the location prior G(mu) as a Dirichlet process itself. This induces a Dirichlet mixture of mixtures of normals, whose analysis may be developed using Gibbs sampling techniques. We discuss these models and their analysis, and illustrate them in the context of neurological response analysis.

Animals↗

Deactivation mechanism of the green fluorescent chromophore.

We report time-resolved fluorescence data for the anion of p-hydroxybenzylidene dimethylimidazolinone (p-HBDI), a model chromophore of the green fluorescence protein, in viscous glycerol-water mixtures over a range of temperatures, T. The markedly nonexponential decay of the excited electronic state is interpreted with the aid of an inhomogeneous model possessing a Gaussian coordinate-dependent sink term. A nonlinear least-squares fitting routine enables us to achieve quantitative fits by adjusting a single activation parameter, which is found to depend linearly on 1/T. We derive an analytic expression for the absolute quantum yield, which is compared with the integrated steady-state fluorescence spectra. The microscopic origins of the model are discussed in terms of two-dimensional dynamics, coupling the phenyl-ring rotation to a swinging mode that brings this flexible molecule to the proximity of a conical intersection on its multidimensional potential energy surface.

Fluorescence↗

Comparison of machine learning and traditional classifiers in glaucoma diagnosis.

Glaucoma is a progressive optic neuropathy with characteristic structural changes in the optic nerve head reflected in the visual field. The visual-field sensitivity test is commonly used in a clinical setting to evaluate glaucoma. Standard automated perimetry (SAP) is a common computerized visual-field test whose output is amenable to machine learning. We compared the performance of a number of machine learning algorithms with STATPAC indexes mean deviation, pattern standard deviation, and corrected pattern standard deviation. The machine learning algorithms studied included multilayer perceptron (MLP), support vector machine (SVM), and linear (LDA) and quadratic discriminant analysis (QDA), Parzen window, mixture of Gaussian (MOG), and mixture of generalized Gaussian (MGG). MLP and SVM are classifiers that work directly on the decision boundary and fall under the discriminative paradigm. Generative classifiers, which first model the data probability density and then perform classification via Bayes' rule, usually give deeper insight into the structure of the data space. We have applied MOG, MGG, LDA, QDA, and Parzen window to the classification of glaucoma from SAP. Performance of the various classifiers was compared by the areas under their receiver operating characteristic curves and by sensitivities (true-positive rates) at chosen specificities (true-negative rates). The machine-learning-type classifiers showed improved performance over the best indexes from STATPAC. Forward-selection and backward-elimination methodology further improved the classification rate and also has the potential to reduce testing time by diminishing the number of visual-field location measurements.

Artificial Intelligence↗

Stability of inverse bicontinuous cubic phases in lipid-water mixtures

We investigate the stability of seven inverse bicontinuous cubic phases [ G, D, P, C(P), S, I-WP, F-RD] in lipid-water mixtures based on a curvature model of membranes. Lipid monolayers are described by parallel surfaces to triply periodic minimal surfaces. The phase behavior is determined by the distribution of the Gaussian curvature on the minimal surface and the porosity of each structure. Only G, D, and P are found to be stable, and to coexist along a triple line. The calculated phase diagram agrees very well with experimental results for 2:1 lauric acid/DLPC.

Journal Article↗

A quantitative and qualitative description of electromyographic linear envelopes for synergy analysis.

The muscular synergy patterns of human locomotion can be described by the phasic activity of electromyographic linear envelopes (LE) and the interphasic spatio-temporal relations. To represent the phasic activity, the LE is modeled as the summation of Gaussian pulses of various lengths. The parameters of interest are the temporal features: time, duration, and amplitude of the phases of activity. A maximum likelihood approach to the parameter estimation for a mixture of normal distributions is adopted for extracting the temporal features. Based on the derived temporal features, a set of relational descriptors can be defined to describe the spatio-temporal relations between the multichannel phasic activities. The strength of this approach is not only that the phasic activity of LE can be quantitatively represented accurately, but also that the resulting synergy patterns can be easily interpreted by observers.

Algorithms↗

Hadamard conjugations and modeling sequence evolution with unequal rates across sites.

This paper considers the many different distributions that may approximate the distribution of site rates in DNA sequences and shows how the Hadamard conjugation may be modified to take these into account. This is done for both 2-state and 4-state data. Distributions which give simple closed forms include the gamma (gamma) distribution, the inverse Gaussian distribution (which is similar to the lognormal), and a mixture of either of these with a proportion of sites which cannot change (invariant sites). It is seen that the tail of a distribution can have major effects upon the coefficient of variation of site rates. Because the Hadamard conjugation can be used to either correct data or predict the data given the model (i.e., the likelihood of site patterns), light is shed on properties of maximum likelihood tree selection with unequal site rates. Analysis of rRNA shows how unequal rates across sites can change the optimal tree. Maximum likelihood analysis also shows that distinct distributions fit each data set, with the gamma often not being the best. Analyzing both these data and a long stretch of primate mtDNA reveals evidence of many "hidden" multiple substitutions, while signals not corresponding to the preferred biological tree generally decrease an unequal rates are allowed for. Last, we discuss the expected behavior of sequences evolving by models where stabilizing selection alone explains unequal site rates. Such models do not explain "synapomorphies" or informative changes in ancient molecules, because while stabilizing selection can vastly decrease change at a site, it will also vastly accelerate back-substitution (leaving only a covarion model to explain old synapomorphies). When and why models allowing a continuous distribution of site rates (e.g., gamma) will approximate covarion evolution requires further study.

Algorithms↗

Canonical functions for dispersal-induced synchrony.

Two processes are universally recognized for inducing spatial synchrony in abundance: dispersal and correlated environmental stochasticity. In the present study we seek the expected relationship between synchrony and distance in populations that are synchronized by density-independent dispersal. In the absence of dispersal, synchrony among populations with simple dynamics has been shown to echo the correlation in the environment. We ask what functional form we may expect between synchrony and distance when dispersal is the synchronizing agent. We formulate a continuous-space, continuous-time model that explicitly represents the time evolution of the spatial covariance as a function of spatial distance. Solving this model gives us two simple canonical functions for dispersal-induced covariance in spatially extended populations. If dispersal is rare relative to birth and death, then covariances between nearby points will follow the dispersal distance distribution. At long distances, however, the covariance tails off according to exponential or Bessel functions (depending on whether the population moves in one or two dimensions). If dispersal is common, then the covariances will follow the mixture distribution that is approximately Gaussian around the origin and with an exponential or Bessel tail. The latter mixture results regardless of the original dispersal distance distribution. There are hence two canonical functions for dispersal-induced synchrony

Animals↗

Fusion of Hidden Markov Random Field models and its Bayesian estimation.

In this paper, we present a Hidden Markov Random Field (HMRF) data-fusion model. The proposed model is applied to the segmentation of natural images based on the fusion of colors and textons into Julesz ensembles. The corresponding Exploration/ Selection/Estimation (ESE) procedure for the estimation of the parameters is presented. This method achieves the estimation of the parameters of the Gaussian kernels, the mixture proportions, the region labels, the number of regions, and the Markov hyper-parameter. Meanwhile, we present a new proof of the asymptotic convergence of the ESE procedure, based on original finite time bounds for the rate of convergence.

Algorithms↗

Supervised cluster analysis for microarray data based on multivariate Gaussian mixture.

MOTIVATION: Grouping genes having similar expression patterns is called gene clustering, which has been proved to be a useful tool for extracting underlying biological information of gene expression data. Many clustering procedures have shown success in microarray gene clustering; most of them belong to the family of heuristic clustering algorithms. Model-based algorithms are alternative clustering algorithms, which are based on the assumption that the whole set of microarray data is a finite mixture of a certain type of distributions with different parameters. Application of the model-based algorithms to unsupervised clustering has been reported. Here, for the first time, we demonstrated the use of the model-based algorithm in supervised clustering of microarray data. RESULTS: We applied the proposed methods to real gene expression data and simulated data. We showed that the supervised model-based algorithm is superior over the unsupervised method and the support vector machines (SVM) method. AVAILABILITY: The program written in the SAS language implementing methods I-III in this report is available upon request. The software of SVMs is available in the website http://svm.sdsc.edu/cgi-bin/nph-SVMsubmit.cgi

Algorithms↗

General time-reversible distances with unequal rates across sites: mixing gamma and inverse Gaussian distributions with invariant sites.

A series of new results useful to the study of DNA sequences using Markov models of substitution are presented with proofs. General time-reversible distances can be extended to accommodate any fixed distribution of rates across sites by replacing the logarithmic function of a matrix with the inverse of a moment generating function. Estimators are presented assuming a gamma distribution, the inverse Gaussian distribution, or a mixture of either of these with invariant sites. Also considered are the different ways invariant sites may be removed and how these differences may affect estimated distances. Through collaboration, we implemented these distances into PAUP in 1994. The variance of these new distances is approximated via the delta method. It is also shown how to predict the divergence expected for a pair of sequences given a rate matrix and a distribution of rates across sites, allowing iterated ML estimates of distances under any reversible model. A simple test of whether a rate matrix is time reversible is also presented. These new methods are used to estimate the divergence time of humans and chimps from mtDNA sequence data. These analyses support suggestions that the human lineage has an enhanced transition rate relative to other hominoids. These studies also show that transversion distances differ substantially from the overall distances which are dominated by transitions. Transversions alone apparently suggest a very recent divergence time for humans versus chimps and/or a very old (> 16 myr) divergence time for humans versus orangutans. This work illustrates graphically ways to interpret the reliability of distance-based transformations, using the corrected transition to transversion ratio returned for pairs of sequences which are successively more diverged.

Animals↗