Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Gaussian mixture model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

General time-reversible distances with unequal rates across sites: mixing gamma and inverse Gaussian distributions with invariant sites.

A series of new results useful to the study of DNA sequences using Markov models of substitution are presented with proofs. General time-reversible distances can be extended to accommodate any fixed distribution of rates across sites by replacing the logarithmic function of a matrix with the inverse of a moment generating function. Estimators are presented assuming a gamma distribution, the inverse Gaussian distribution, or a mixture of either of these with invariant sites. Also considered are the different ways invariant sites may be removed and how these differences may affect estimated distances. Through collaboration, we implemented these distances into PAUP in 1994. The variance of these new distances is approximated via the delta method. It is also shown how to predict the divergence expected for a pair of sequences given a rate matrix and a distribution of rates across sites, allowing iterated ML estimates of distances under any reversible model. A simple test of whether a rate matrix is time reversible is also presented. These new methods are used to estimate the divergence time of humans and chimps from mtDNA sequence data. These analyses support suggestions that the human lineage has an enhanced transition rate relative to other hominoids. These studies also show that transversion distances differ substantially from the overall distances which are dominated by transitions. Transversions alone apparently suggest a very recent divergence time for humans versus chimps and/or a very old (> 16 myr) divergence time for humans versus orangutans. This work illustrates graphically ways to interpret the reliability of distance-based transformations, using the corrected transition to transversion ratio returned for pairs of sequences which are successively more diverged.

Animals↗

Analysis of overdispersed count data by mixtures of Poisson variables and Poisson processes.

Count data often show overdispersion compared to the Poisson distribution. Overdispersion is typically modeled by a random effect for the mean, based on the gamma distribution, leading to the negative binomial distribution for the count. This paper considers a larger family of mixture distributions, including the inverse Gaussian mixture distribution. It is demonstrated that it gives a significantly better fit for a data set on the frequency of epileptic seizures. The same approach can be used to generate counting processes from Poisson processes, where the rate or the time is random. A random rate corresponds to variation between patients, whereas a random time corresponds to variation within patients.

Anticonvulsants↗

Atmospheric dispersion and deposition of 131I released from the Hanford Site.

Approximately 2.6 x 10(4) TBq (700,000 Ci) of 131I were released to the air from reactor fuel processing plants on the Hanford Site in southcentral Washington State from December 1944 through December 1949. The Hanford Environmental Dose Reconstruction Project developed a suite of codes to estimate the doses that might have resulted from these releases. The Regional Atmospheric Transport Code for Hanford Emission Tracking (RATCHET) computer code is part of this suite. The RATCHET code implements a Lagrangian-trajectory, Gaussian-puff dispersion model that uses hourly meteorological and release rate data to estimate daily time-integrated air concentrations and surface contamination for use in dose estimates. In this model, iodine is treated as a mixture of three species (inorganic gases, organic gases, and particles). Model deposition parameters are functions of the mixture and meteorological conditions. A resistance model is used to calculate dry deposition velocities. Equilibrium between concentrations in the precipitation and the air near the ground is assumed in calculating wet deposition of gases, and irreversible washout of the particles is assumed. RATCHET explicitly treats the uncertainties in model parameters and meteorological conditions. Uncertainties in 131I release rates and partitioning among the nominal species are treated by varying model input. The results of 100 model runs for December 1944 through December 1949 indicate that monthly average air concentrations and deposition have uncertainties ranging from a factor of two near the center of the time-integrated plume to more than an order of magnitude near the edge. These results indicate that approximately 10% of the 131I released to the atmosphere decayed during transit in the study area, approximately 56% was deposited within the study area, and the remaining 34% was transported out of the study area while still in the air.

Air Pollutants, Radioactive↗

Steady-state concentration distribution of ampholytes in isoelectric focusing in a linear immobilized pH gradient.

The equation of balance between electrophoretic and diffusional mass flows in the steady state of isoelectric focusing is analyzed. To create the pH gradient, a model system composed of only one Immobiline is used. The solution is found for the case of a small sample concentration, but without assumptions about linear "focusing force" and constant (linear) conductivity profiles. The effect of sample concentration on the final concentration profile is also evaluated. At high sample concentrations, it is demonstrated that the steady-state distribution is essentially non-Gaussian and results in a considerably lower concentration maximum as compared with low sample levels.

Ampholyte Mixtures↗

Tests for genetic linkage and homogeneity.

This report concerns likelihood ratio tests in a heterogeneous model for linkage, where the recombination fraction has a binomial mixture distribution with an unknown proportion of unlinked families. We consider families of unequal sizes with known or unknown phase data. In both cases, the limit distributions of the linkage test statistics are a mixture of a mass at ) and of a X2/1 distribution in equal proportions and homogeneity test statistics tend to the supremum of Gaussian processes. The critical values of the homogeneity tests are simulated, and the power functions of the linkage and homogeneity tests are compared in a simulation study.

Biometry↗

Making the most of a patient's laboratory data: optimisation of signal-to-noise ratio.

All results in laboratory medicine are compared to some reference for interpretation. This reference may be a previous result from the same patient, a reference population--either healthy or diseased, or both--or a decision limit recommended by an expert group. The aim for the medical laboratory is to improve the signal-to-noise ratio by increasing the signal or reducing the noise. This presentation deals with the more general tools for reduction of the noise component, and focuses on biological within-subject variation, reference intervals and decision models. Regarding biological within-subject variation, the estimation of reference change value (RCV) as a yardstick for judging measured differences within the patient over time is an important tool. Here, only type 1 errors are usually applied, but type 2 errors should also be taken into consideration. Moreover, variance homogeneity is assumed for the application of RCV, but this assumption is not always fulfilled, and erroneous interpretations may be introduced. A tool for comparison of different and more complicated algorithms applied to serial measurements is computer simulation (e.g. on data from tumour markers). In order to reduce the noise component from reference intervals, partitioning according to relevant subgroups is a tool, and useful criteria for judging whether subgroups should be combined are reported. Geographical and racial differences may cause different reference distributions (e.g. plasma proteins), but it has been possible to establish common reference intervals for 25 common components in Caucasians in the five Nordic countries. Transformation of data and presentation of accumulated ranked values in rankit plots where Gaussian (or log-Gaussian) distributions show up as straight lines is a valuable tool for interpretation of the distributions and comparison of subgroups. In this way it is often possible to isolate a low-risk group which fits a log-Gaussian distribution. In case of thyroid autoantibodies the distributions look biphasic, even after all possible rule-out criteria have been exhausted, but a composite model makes it possible to extract a reference population from the mixture. The classical decision model is bimodal, reflecting an assumption of two independent but overlapping distributions, with a clear but unknown prevalence for the disease. When a decision limit (cut-off point) is applied, the percentages of false positive and false negative results will be determined, but the underlying prevalence is unchanged. In contrast to this, the unimodal distributions cover a continuum of probabilities for a certain disease (risk) and the decision limit is arbitrarily chosen. Thus, the disease or risk is defined by the measured component. Consequently, the decision limit directly defines the prevalence, and this limit can be different over time and geography as has been the case for cholesterol or glucose.

Journal Article↗

Entropy-based kernel mixture modeling for topographic map formation.

A new information-theoretic learning algorithm for kernel-based topographic map formation is introduced. In the one-dimensional case, the algorithm is aimed at uniformizing the cumulative distribution of the kernel mixture densities by maximizing its differential entropy. A nonparametric differential entropy estimator is used on which normalized gradient ascent is performed. Both differentiable and nondifferentiable kernels are in principle supported, such as Gaussian and rectangular (on/off) kernels. The relation is shown with joint entropy maximization of the kernel outputs. The learning algorithm's performance is assessed and compared with the theoretically optimal performance. A fixed-point rule is derived for the case of heterogeneous kernel mixtures. Finally, an extension of the algorithm to the multidimensional case is suggested.

Algorithms↗

Statistical evaluation of cell kinetic data from DNA flow cytometry (FCM) by the EM algorithm.

Flow cytometric DNA measurements yield the amount of DNA for each of a large number of cells. A DNA histogram normally consists of a mixture of one or more constellations of G0/G1-, S-, G2/M-phase cells, together with internal standards, debris, background noise, and one or more populations of clumped cells. We have modelled typical DNA histograms as a mixed distribution with Gaussian densities for the G0/G1 and G2/M phases, an S-phase density, assumed to be uniform between the G0/G1 and G2/M peaks, observed with a Gaussian error, and with Gaussian densities for standards of chicken and trout red blood cells. The debris is modelled as a truncated exponential distribution, and we also have included a uniform background noise distribution over the whole observation interval. We have explored a new approach for maximum-likelihood analyses of complex DNA histograms by the application of the EM algorithm. This algorithm was used for four observed DNA histograms of varying complexity. Our results show that the algorithm works very well, and it converges to reasonable values for all parameters. In simulations from the estimated models, we have investigated bias, variance, and correlations of the estimates.

Algorithms↗

Investigations of the simian ontogenic switch from fetal to adult hemoglobin at the progenitor cell level.

The ontogenic switch from fetal to adult hemoglobin could result from discontinuous events, such as replacement of fetal erythroid progenitor cells by adult ones, or gradual modulation of the hemoglobin program of a single progenitor cell pool. The former would result in progenitors at midswitch with skewed fractional beta-globin synthesis programs, the latter in a Gaussian distribution. For these studies, we obtained bone marrow from rhesus monkey fetuses at 141-153 d (midswitch). Mononuclear cells were cultured in methyl cellulose with erythropoietin, and single BFU-E-derived colonies were removed and incubated with [3H]leucine. Globin synthesis was examined by gel electrophoresis and fluorography. The beta-globin synthesis pattern of single fetal colonies was skewed, and did not fit a normal distribution. The fetal pattern resembled the pattern of an artificial mixture of fetal and adult progenitors, suggesting that the fetal progenitor pool could contain populations with different beta-globin programs. This non-Gaussian distribution in the progenitors of midswitch fetuses is consistent with a discontinuous model for hemoglobin switching during ontogeny.

Animals↗

Pair dynamics in a glass-forming binary mixture: simulations and theory.

We have carried out molecular dynamics simulations to understand the dynamics of a tagged pair of atoms in a strongly nonideal glass-forming binary Lennard-Jones mixture. Here atom B is smaller than atom A (sigma(BB)=0.88sigma(AA), where sigma(AA) is the molecular diameter of the A particles) and the AB interaction is stronger than that given by Lorentz-Berthelot mixing rule (epsilon(AB)=1.5epsilon(AA), where epsilon(AA) is the interaction energy strength between the A particles). The generalized time-dependent pair distribution function is calculated separately for the three pairs (AA, BB, and AB). The three pairs are found to behave differently. The relative diffusion constants are found to vary in the order D(BB)(R)>D(AB)(R)>D(AA)(R), with D(BB)(R) approximately 2D(AA)(R), showing the importance of the hopping process (B hops much more than A). We introduce a non-Gaussian parameter [alpha(P)(2)(t)] to monitor the relative motion of a pair of atoms and evaluate it for all the three pairs with initial separations chosen to be at the first peak of the corresponding partial radial distribution functions. At intermediate times, significant deviation from the Gaussian behavior of the pair distribution functions is observed with different degrees for the three pairs. A simple mean-field (MF) model, proposed originally by Haan [Phys. Rev. A 20, 2516 (1979)] for one-component liquid, is applied to the case of a binary mixture and compared with the simulation results. While the MF model successfully describes the dynamics of the AA and AB pairs, the agreement for the BB pair is less satisfactory. This is attributed to the large scale anharmonic motions of the B particles in a weak effective potential. Dynamics of the next nearest neighbor pairs is also investigated.

Journal Article↗

Analysis of fluorophore diffusion by continuous distributions of diffusion coefficients: application to photobleaching measurements of multicomponent and anomalous diffusion.

Fluorescence recovery after photobleaching (FRAP) is widely used to measure fluorophore diffusion in artificial solutions and cellular compartments. Two new strategies to analyze FRAP data were investigated theoretically and applied to complex systems with anomalous diffusion or multiple diffusing species: 1) continuous distributions of diffusion coefficients, alpha(D), and 2) time-dependent diffusion coefficients, D(t). A regression procedure utilizing the maximum entropy method was developed to resolve alpha(D) from fluorescence recovery curves, F(t). The recovery of multi-component alpha(D) from simulated F(t) with random noise was demonstrated and limitations of the method were defined. Single narrow Gaussian alpha(D) were recovered for FRAP measurements of thin films of fluorescein and size-fractionated FITC-dextrans and Ficolls, and multi-component alpha(D) were recovered for defined fluorophore mixtures. Single Gaussian alpha(D) were also recovered for solute diffusion in viscous media containing high dextran concentrations. To identify anomalous diffusion from FRAP data, a theory was developed to compute F(t) and alpha(D) for anomalous diffusion models defined by arbitrary nonlinear mean-squared displacement versus time relations. Several characteristic alpha(D) profiles for anomalous diffusion were found, including broad alpha(D) for subdiffusion, and alpha(D) with negative amplitudes for superdiffusion. A method to deduce apparent D(t) from F(t) was also developed and shown to provide useful complementary information to alpha(D). alpha(D) and D(t) were determined from photobleaching measurements of systems with apparent anomalous subdiffusion (nonuniform solution layer) and superdiffusion (moving fluid layer). The results establish a practical strategy to characterize complex diffusive phenomena from photobleaching recovery measurements.

Biophysical Phenomena↗

Lifetime distributions and anisotropy decays of indole fluorescence in cyclohexane/ethanol mixtures by frequency-domain fluorometry.

We used frequency-domain fluorometry to measure intensity and anisotropy decay of indole fluorescence in cyclohexane/ethanol mixtures at 20 degrees C. In 100% cyclohexane or 100% ethanol the intensity decay of indole appears to be a single exponential with decay times of 7.66 and 4.10 ns, respectively. In cyclohexane containing a small percentage of ethanol (up to 10%), we observed increased heterogeneity in intensity decay, resulting in a 10-fold increase in chi 2R for the single-exponential fit, as compared with the double-exponential model. We obtained comparable or better fits using unimodal Lorentzian and Gaussian lifetime distributions (two floating parameters) than for the two-exponential model (three floating parameters). We believe that the distribution of decay times reflects a range of indole solvation states in the dominately nonpolar solutions. This result suggests that a variety of hydrogen-bonding configurations could be one origin of the distributions of decay times observed for tryptophan emission from proteins. We also measured rotational diffusion of indole in cyclohexane, ethanol and its mixtures at 20 degrees C. The picosecond correlation times required that the mean decay times be decreased by acrylamide quenching (in ethanol) or energy transfer (in cyclohexane). In ethanol we observed nearly isotropic rotation of indole; in cyclohexane we obtained two correlation times of 17 and 73 ps. The shorter correlation time in cyclohexane appears to be due to the slip boundary condition, which was found to be progressively eliminated by small percentages of ethanol. Hence, hydrogen-bonding interactions appear to have a substantial effect on the rotational dynamics of indole.

Cyclohexanes↗

Surface curvatures of trabecular bone microarchitecture.

Microstructure of trabecular bone has been examined with a particular emphasis on surface curvatures in two-phase (trabecular and intertrabecular space- i.e., marrow space) structures. Three trabecular bone samples, quantified as "plate-like," "rod-like," and a mixture of these two structural elements according to the structure model index (SMI), were subjected to analysis based on (differential) geometry. A correspondence between the SMI and the mean curvature was found. A method to measure surface curvatures is proposed. The gaussian curvatures averaged over the surfaces for the three analyzed bone structures were all found to be negative, demonstrating their surfaces to be, on average, hyperbolic. In addition, the Euler-Poincaré characteristics and the genus, both characterizing topological features of bone connectivity, were estimated from integral gaussian curvature (Gauss-Bonnet theorem). The three bone microstructures were found to be topologically analogous to spheres with one to three handles.

Bone and Bones↗

Skin segmentation using color pixel classification: analysis and comparison.

This paper presents a study of three important issues of the color pixel classification approach to skin segmentation: color representation, color quantization, and classification algorithm. Our analysis of several representative color spaces using the Bayesian classifier with the histogram technique shows that skin segmentation based on color pixel classification is largely unaffected by the choice of the color space. However, segmentation performance degrades when only chrominance channels are used in classification. Furthermore, we find that color quantization can be as low as 64 bins per channel, although higher histogram sizes give better segmentation performance. The Bayesian classifier with the histogram technique and the multilayer perceptron classifier are found to perform better compared to other tested classifiers, including three piecewise linear classifiers, three unimodal Gaussian classifiers, and a Gaussian mixture classifier.

Algorithms↗

Peak deconvolution in one-dimensional chromatography using a two-way data approach.

A deconvolution methodology for overlapped chromatographic signals is proposed. Several single-wavelength chromatograms of binary mixtures, obtained in different runs at diverse concentration ratios of the individual components, were simultaneously processed (multi-batch approach), after being arranged as two-way data. The chromatograms were modelled as linear combinations of forced peak profiles according to a polynomially modified Gaussian equation. The fitting was performed with a previously reported hybrid genetic algorithm with local search, leaving all model parameters free. The approach yielded more accurate solutions than those found when each experimental chromatogram was fitted independently to the peak model (single-batch approach). The improvement was especially significant for those chromatograms where the peaks were severely affected by the tails of the preceding compounds. Peak shifts among chromatograms, which are a usual source of non-bilinearity, were modelled in a continuous domain instead of in a discrete way, which avoided some drawbacks associated with latent variable methods. An experimental design involving simulated chromatograms was applied to check the method performance. Five main factors affecting the deconvolution were examined: concentration pattern, chromatographic resolution, number of batches and replicates, and noise level, which were evaluated using first- and second-order figures of merit. The method was also tested on three real samples containing compounds showing different overlap. Four multi-batch deconvolution methods were considered differing in the nature of the processed information and kind of peak matching among chromatograms. In all cases, the multi-batch deconvolution yielded better performance than the single-batch approach.

Chromatography↗

[Evaluation of partial volume effect in quantitative measurement of regional cerebral blood flow using positron emission tomography].

Effects of limited spatial resolution of the positron emission tomography (PET) scanner on the quantitative measurement of regional cerebral blood flow (rCBF) was investigated for various tracer kinetic models with use of 15O labeled water and PET. Using a numerical brain phantom consisting of a gray matter, white matter and cerebrospinal fluid components, dynamical tracer distribution images were calculated for the H215O bolus injection and for the C15O2 gas inhalation protocols. The tracer distribution images were convoluted with a 2 dimensional gaussian function with full-width at half maximum (FWHM) of 4, 7, 12 mm to simulate a limited spatial resolution of the PET scanner, and rCBF images were calculated according to some kinetic models. Smoothing the tracer distribution images caused a heterogeneous structure (tissue mixture) in a given volume element. rCBF values calculated by models with use of 15O-water and PET were found to provide rCBF values that were systematically underestimated compared with those obtained by the microsphere model for a mixed tissue region. Moreover, the magnitude of the underestimation was shown to be highly dependent on the tracer kinetic models employed, those errors for mixed tissue of gray and white matter were 20% on steady state, 9% on autoradiography and on weighted integration method, and 2% on non-linear least squares fitting, compared with microsphere model. More errors observed by steady state method and autoradiography method happened for tissue mixture consisting gray matter, white matter and cerebrospinal fluid components. It is important to take into account for difference of the partial volume effect for each models in calculated rCBF.

Brain↗

Mixture distributions in psychiatric research.

This paper describes the application of Gaussian mixture distributions to biological marker research in psychiatry. Mixtures of univariate and multivariate normal distributions can be used to determine if diagnostically similar psychiatric patients belong to biologically distinct subpopulations. The resulting biological subtypes may be important in understanding the etiology of psychiatric disorders. The general model and estimation procedure are described (EM algorithm; Dempster, Laird and Rubin 1977). The method is illustrated using two examples of biological data: (1) red cell membranes and monoamine oxidase activity data in normal individuals having no family history of psychiatric illness, the first-degree relatives of bipolar depressed patients and a heterogeneous patient population; and (2) smooth pursuit eye movements that classify relatives of schizophrenics, nonschizophrenics and normal controls into biologically distinct populations.

Bipolar Disorder↗

Quantifying biological specificity: the statistical mechanics of molecular recognition.

The Random Energy Model of statistical physics is applied to the problem of the specificity of recognition between two biological (macro)molecules forming a non-covalent complex. In this model, the native mode of association is separated by an energy gap from a large body of non-native modes. Whereas the native mode is unique, the non-native modes form an energy spectrum which is approximated by a gaussian distribution. Specificity can then be estimated by writing the partition function and calculating the ratio r of non-native to native modes at thermodynamic equilibrium. We examine three situations: (i) recognition in the absence of a competitor; (ii) recognition in the presence of a competing ligand; (iii) recognition in a heterogeneous mixture. We derive the dependence of the ratio r on temperature and on the concentration of competing ligands, and we estimate the effect of a local perturbation such as can result from a point mutation. Cases (i) and (iii) are modeled by docking experiments in the computer. In case (iii), which is representative of a wide variety of biological situations, we show that increasing the heterogeneity of a mixture affects the specificity of recognition, even when the concentration of competing species is kept constant.

Binding Sites↗