Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

In vitro dissolution profile comparison--statistics and analysis of the similarity factor, f2.

PURPOSE: To describe the properties of the similarity factor (f2) as a measure for assessing the similarity of two dissolution profiles. Discuss the statistical properties of the estimate based on sample means. METHODS: The f2 metrics and the decision rule is evaluated using examples of dissolution profiles. The confidence interval is calculated using bootstrapping method. The bias of the estimate using sample mean dissolution is evaluated. RESULTS: 1. f2 values were found to be sensitive to number of sample points, after the dissolution plateau has been reached. 2. The statistical evaluation of f2 could be made using 90% confidence interval approach. 3. The statistical distribution of f2 metrics could be simulated using 'Bootstrap' method. A relatively robust distribution could be obtained after more than 500 'Bootstraps'. 4. A statistical 'bias correction' was found to reduce the bias. CONCLUSIONS: The similarity factor f2 is a simple measure for the comparison of two dissolution profiles. But the commonly used similarity factor estimate f2 is a biased and conservative estimate of f2. The bootstrap approach is a useful tool to simulate the confidence interval.

Chemistry, Pharmaceutical↗

Optimizing statistical Shake-and-Bake for Se-atom substructure determination.

A novel statistical approach to the phase problem in X-ray crystallography was introduced in a recent paper [Xu & Hauptman (2004), Acta Cryst. A60, 153-157]. In this approach, a new minimal function based on the statistical distribution of structure-invariant values serves as the foundation of an optimization procedure called statistical Shake-and-Bake. Favorable application of this procedure to Se-atom substructure determination depends on the choice of the statistical interval over which the function is defined. The effects of interval variation have been studied for 19 Se-atom substructures ranging in size from five to 70 Se atoms in the asymmetric unit and the results have shown an overall improvement in success rate relative to traditional Shake-and-Bake. Statistical Shake-and-Bake is being incorporated as the default optimization procedure in newly distributed versions of the SnB and BnP computer programs.

Crystallography, X-Ray↗

Advanced statistics: bootstrapping confidence intervals for statistics with "difficult" distributions.

The use of confidence intervals in reporting results of research has increased dramatically and is now required or highly recommended by editors of many scientific journals. Many resources describe methods for computing confidence intervals for statistics with mathematically simple distributions. Computing confidence intervals for descriptive statistics with distributions that are difficult to represent mathematically is more challenging. The bootstrap is a computationally intensive statistical technique that allows the researcher to make inferences from data without making strong distributional assumptions about the data or the statistic being calculated. This allows the researcher to estimate confidence intervals for statistics that do not have simple sampling distributions (e.g., the median). The purposes of this article are to describe the concept of bootstrapping, to demonstrate how to estimate confidence intervals for the median and the Spearman rank correlation coefficient for non-normally-distributed data from a recent clinical study using two commonly used statistical software packages (SAS and Stata), and to discuss specific limitations of the bootstrap.

Confidence Intervals↗

Nephrotoxicity screening in rats; general approach and establishment of test criteria.

The concept of a nephrotoxicity screening test that is based on quantitative assessment of urine collected under standardized conditions for 15.5 h is presented. One to eight urine collections were performed in large numbers of untreated female Sprague-Dawley rats. Normal values for water consumption, urine volume, pH, and excretion of protein, gamma-glutamyltranspeptidase, malate dehydrogenase, electrolytes, glucose, amino acids, leukocytes, erythrocytes, epithelia, unspecified cells and cylinders were determined. Test criteria were established based on the statistical distribution of these measurements. In rats repeatedly placed in metabolism cages, a statistically significant decrease in leukocyte excretion and an increase in excretion of epithelia and unspecified cells were observed. All other variables did not change with time.

Animals↗

A non-invasive method for in situ quantification of subpopulation behaviour in mixed cell culture.

Ongoing advances in quantitative molecular- and cellular-biology highlight the need for correspondingly quantitative methods in tissue-biology, in which the presence and activity of specific cell-subpopulations can be assessed in situ. However, many experimental techniques disturb the natural tissue balance, making it difficult to draw realistic conclusions concerning in situ cell behaviour. In this study, we present a widely applicable and minimally invasive method which combines fluorescence cell labelling, retrospective image analysis and mathematical data processing to detect the presence and activity of cell subpopulations, using adhesion patterns in STRO-1 immunoselected human mesenchymal populations and the homogeneous osteoblast-like MG63 continuous cell line as an illustration. Adhesion is considered on tissue culture plastic and fibronectin surfaces, using cell area as a readily obtainable and individual cell specific measure of spreading. The underlying statistical distributions of cell areas are investigated and mappings between distributions are examined using a combination of graphical and non-parametric statistical methods. We show that activity can be quantified in subpopulations as small as 1% by cell number, and outline behaviour of significant subpopulations in both STRO-1+/- fractions. This method has considerable potential to understand in situ cell behaviour and thus has wide applicability, for example in developmental biology and tissue engineering.

Cell Line↗

Wave scattering through classically chaotic cavities in the presence of absorption: An information-theoretic model

We propose an information-theoretic model for the transport of waves through a chaotic cavity in the presence of absorption. The entropy of the S-matrix statistical distribution is maximized, with the constraint =alphan: n is the dimensionality of S, and 0</=alpha</=1, alpha=0(1) meaning complete (no) absorption. For strong absorption our result agrees with a number of analytical calculations already given in the literature. In that limit, the distribution of the individual (angular) transmission and reflection coefficients becomes exponential (Rayleigh statistics), even for n=1. For n>>1 Rayleigh statistics is attained even with no absorption; here, we extend the study to alpha<1. The model is compared with random-matrix-theory numerical simulations: it describes the problem very well for strong absorption, but fails for moderate and weak absorptions. Thus, in the latter regime, some important physical constraint is missing in the construction of the model.

Journal Article↗

Reduction of noise-induced streak artifacts in X-ray computed tomography through spline-based penalized-likelihood sinogram smoothing.

We present a statistically principled sinogram smoothing approach for X-ray computed tomography (CT) with the intent of reducing noise-induced streak artifacts. These artifacts arise in CT when some subset of the transmission measurements capture relatively few photons because of high attenuation along the measurement lines. Attempts to reduce these artifacts have focused on the use of adaptive filters that strive to tailor the degree of smoothing to the local noise levels in the measurements. While these approaches involve loose consideration of the measurement statistics to determine smoothing levels, they do not explicitly model the statistical distributions of the measurement data. In this paper, we present an explicitly statistical approach to sinogram smoothing in the presence of photon-starved measurements. It is an extension of a nonparametric sinogram smoothing approach using penalized Poisson-likelihood functions that we have previously developed for emission tomography. Because the approach explicitly models the data statistics, it is naturally adaptive--it will smooth more variable measurements more heavily than it does less variable measurements. We find that it significantly reduces streak artifacts and noise levels without comprising image resolution.

Algorithms↗

Regional septal dysfunction in a three-dimensional computational model of focal myofiber disarray.

MLC2v/ras transgenic mice display a phenotype characteristic of hypertrophic cardiomyopathy, with septal hypertrophy and focal myocyte disarray. Experimental measurements of septal wall mechanics in ras transgenic mice have previously shown that regions of myocyte disarray have reduced principal systolic shortening, torsional systolic shear, and sarcomere length. To investigate the mechanisms of this regional dysfunction, a three-dimensional prolate spheroidal finite-element model was used to simulate filling and ejection in the hypertrophied mouse left ventricle with septal disarray. Focally disarrayed septal myocardium was modeled by randomly distributed three-dimensional regions of altered material properties based on measured statistical distributions of muscle fiber angular dispersion. Material properties in disarrayed regions were modeled by decreased systolic anisotropy derived from increased fiber angle dispersion and decreased systolic tension development associated with reduced sarcomere lengths. Compared with measurements in ras transgenic mice, the model showed similar heterogeneity of septal systolic strain with the largest reductions in principal shortening and torsional shear in regions of greatest disarray. Average systolic principal shortening on the right ventricular septal surface of the model was -0.114 for normal regions and -0.065 for disarrayed regions; for torsional shear, these values were 0.047 and 0.019, respectively. These model results suggest that regional dysfunction in ras transgenic mice may be explained in part by the observed structural defects, including myofiber dispersion and reduced sarcomere length, which contributed about equally to predicted dysfunction in the disarrayed myocardium.

Animals↗

Centile charts I: new method of assessment for univariate reference intervals.

BACKGROUND: We introduce a new criterion, the percentile inclusion probability, for comparing methods for calculating reference intervals. The criterion is compared with a previously published measure of reliability suggested by Linnet (Linnet K. Clin Chem 1987;33:381-6), the ratio of the width of the confidence interval for the percentile to that of the reference interval. METHODS: Data were simulated from a range of theoretical statistical distributions representing the shapes of data sets encountered in clinical investigations. The two-stage transformation of the data to a gaussian distribution recommended by the IFCC was compared with a nonparametric approach. RESULTS: The percentile inclusion probability criterion identified that the parametric approach is in some cases seriously affected by bias. Using different parametric models, we compared nonparametric and parametric methods for two sets of clinical data and showed that the parametric approach is susceptible to model choice. CONCLUSIONS: Sample sizes significantly greater than those currently recommended are required to establish reference intervals, regardless of whether parametric or nonparametric methods are used. Parametric methods are preferable when the data are truly gaussian, but are only marginally better than nonparametric methods when data transformation is needed to achieve a gaussian shape.

Birth Weight↗

Statistical motor number estimation assuming a binomial distribution.

The statistical method of motor unit number estimation (MUNE) uses the natural stochastic variation in a muscle's compound response to electrical stimulation to obtain an estimate of the number of recruitable motor units. The current method assumes that this variation follows a Poisson distribution. We present an alternative that instead assumes a binomial distribution. Results of computer simulations and of a pilot study on 19 healthy subjects showed that the binomial MUNE values are considerably higher than those of the Poisson method, and in better agreement with the results of other MUNE techniques. In addition, simulation results predict that the performance in patients with severe motor unit loss will be better for the binomial than Poisson method. The adapted method remains closer to physiology, because it can accommodate the increase in activation probability that results from rising stimulus intensity. It does not need recording windows as used with the Poisson method, and is therefore less user-dependent and more objective and quicker in its operation. For these reasons, we believe that the proposed modifications may lead to significant improvements in the statistical MUNE technique.

Adult↗

Preventive distinction of patients with primary or secondary hypertension by discriminant analysis of chronobiologic parameters estimated on 24-hour blood pressure patterns.

This investigation deals with a statistical probatory that patients with primary (PH) or secondary (SH) hypertension may be correctly diagnosed by a discriminant analysis of the chronobiologic characteristics computed on the 24-hour blood pressure (BP) patterns. The methodology concerning non-invasive 24-h BP monitoring, chronobiologic analysis and the discrimination process is detailed. Substantial dissimilarities were found in the statistical distribution for systolic and diastolic BP rhythmometric parameters (mesor, amplitude and acrophase) by a retrospective assessment of two groups, consisting of 54 patients with PH and 16 patients with SH. The group-related distribution for rhythmometric parameters was found to be significantly different to generate a statistically significant intergroup discriminatory boundary. The discriminant analysis correctly diagnosed patients with PH and SH in a percentage of about 91% and 63%, respectively. The high incidence of success is convincing that the combination of 24-h BP monitoring/chronobiologic analysis/discrimination process cna be a practical tool for confidently selecting patients with a presumable PH or SH.

Adolescent↗

Surface vibrational structure at alkane liquid/vapor interfaces.

Broadband vibrational sum frequency spectroscopy (VSFS) has been used to examine the surface structure of alkane liquid/vapor interfaces. The alkanes range in length from n-nonane (C(9)H(20)) to n-heptadecane (C(17)H(36)), and all liquids except heptadecane are studied at temperatures well above their bulk (and surface) freezing temperatures. Intensities of vibrational bands in the CH stretching region acquired under different polarization conditions show systematic, chain length dependent changes. Data provide clear evidence of methyl group segregation at the liquid/vapor interface, but two different models of alkane chain structure can predict chain length dependent changes in band intensities. Each model leads to a different interpretation of the extent to which different chain segments contribute to the anisotropic interfacial region. One model postulates that changes in vibrational band intensities arise solely from a reduced surface coverage of methyl groups as alkane chain length increases. The additional methylene groups at the surface must be randomly distributed and make no net contribution to the observed VSF spectra. The second model considers a simple statistical distribution of methyl and methylene groups populating a three dimensional, interfacial lattice. This statistical picture implies that the VSF signal arises from a region extending several functional groups into the bulk liquid, and that the growing fraction of methylene groups in longer chain alkanes bears responsibility for the observed spectral changes. The data and resulting interpretations provide clear benchmarks for emerging theories of molecular structure and organization at liquid surfaces, especially for liquids lacking strong polar ordering.

Journal Article↗

Monomer composition and sequence of alginates from Pseudomonas aeruginosa.

Alginates from four strains of Pseudomonas aeruginosa, one mucoid strain isolated from a technical water system, one strain isolated from a patient with cystic fibrosis and two mutants of this strain with a defect which affects the O-acetylation of the extracellular alginate, have been isolated and analysed for monomer composition and sequence by 13C-nuclear magnetic resonance (NMR) spectroscopy. The detected contributions of different monomer triplets (triads) were compared with values expected from a statistical chain constitution based on the given monomer ratio. While a typical algal alginate presents a nearly statistical distribution of uronic acids in the polymer chain, a strong deviation from the statistical arrangement of mannuronate (M) and guluronate (G) was found in the alginate of the mucoid strains of P. aeruginosa, being most expressed for the triad MMM. This feature is partially lost in the alginate from the mutant strains, indicating that the O-acetylation is linked to a mechanism which takes influence on the chain sequence. The strong preference for MG-pairs in the parent strain of P. aeruginosa may be connected to a stronger binding of cations in the MG-vicinity.

Acetylation↗

The use of Hasse diagrams as a potential approach for inverse QSAR.

Quantitative structure-activity relationships are often based on standard multidimensional statistical analyses and sophisticated local and global molecular descriptors. Here, the aim is to develop a tool helpful to define a molecule or a class of molecules which fulfills pre-described properties, i.e., an Inverse QSAR approach. If highly sophisticated descriptors are used in QSAR, the structure and then the synthesis recipe may be hard to derive. Thus, descriptors, from which the synthesis recipe can be easily derived, seem appropriate to be included within this study. However, if descriptors simple enough to be useful for defining syntheses recipes of chemicals were used, the accuracy of a numeric expression may fail. This paper suggests a method, based on very simple elements of the theory of partially ordered sets, to find a qualitative basis for the relationship between such fairly simple descriptors on the one side and a series of ecotoxicological properties, on the other side. The partial order ranking method assumes neither linearity nor certain statistical distribution properties. Therefore the method may be more general compared to many standard statistical techniques. A series of chlorinated aliphatic compounds has been used as an illustrative example and a comparison with more sophisticated descriptors derived from quantum chemistry and graph theory is given. Among the results, it was disclosed that only for algae lethal concentration, as one of the four ecotoxicological properties, the synthesis specific predictors seem to be good estimators. For all other ecotoxicological properties quantum chemical descriptors appear as the more suitable estimators.

Ecosystem↗

Dangers and problems in calculating coefficients of a sum of exponential functions.

The use of the computer in making calculations has led increasingly to the estimation of parameters of exponential functions from point experimental measurements. This involves the use of techniques such as logarithmic transformations and criteria (especially that of the least squares), of which the use is not always justified. Radioactivity measurements raise special problems as they are subject to a statistical distribution of the Poisson type. The authors propose a method based on the statistical criterion of "maximum likelihood" which permits tests of the number of exponentials from point measurements (in practice, this method is valid for one or two exponentials), and also the calculation of parameters more satisfactorily than customary methods.

Mathematics↗

[Distribution of 5-methylcytosine in phage lambda genome methylated by DNA methylase Eco RII].

The distribution of 5-methylcytosine in Eco RI-Bam HI fragments of phage lambda DNA in vitro methylated by Eco RII methylase has been studied. The general picture of distribution of methylated sites in phage lambda DNA is slightly different from the statistical distribution. However, the sites have been found, where the distribution of 5-methylcytosine is not accidental. A complete absence of 5-methylcytosine in the J-fragment, a genome lambda area essential for site-specific recombination, has been found. The absence of Eco RII is supposed to be the best protection of this area of phage genome from the increased mutagenesis, characteristic for nucleotide sequences methylated by DNA-methylated Eco RII and Eco RII type.

5-Methylcytosine↗

Feature selection and classifier performance in computer-aided diagnosis: the effect of finite sample size.

In computer-aided diagnosis (CAD), a frequently used approach for distinguishing normal and abnormal cases is first to extract potentially useful features for the classification task. Effective features are then selected from this entire pool of available features. Finally, a classifier is designed using the selected features. In this study, we investigated the effect of finite sample size on classification accuracy when classifier design involves stepwise feature selection in linear discriminant analysis, which is the most commonly used feature selection algorithm for linear classifiers. The feature selection and the classifier coefficient estimation steps were considered to be cascading stages in the classifier design process. We compared the performance of the classifier when feature selection was performed on the design samples alone and on the entire set of available samples, which consisted of design and test samples. The area Az under the receiver operating characteristic curve was used as our performance measure. After linear classifier coefficient estimation using the design samples, we studied the hold-out and resubstitution performance estimates. The two classes were assumed to have multidimensional Gaussian distributions, with a large number of features available for feature selection. We investigated the dependence of feature selection performance on the covariance matrices and means for the two classes, and examined the effects of sample size, number of available features, and parameters of stepwise feature selection on classifier bias. Our results indicated that the resubstitution estimate was always optimistically biased, except in cases where the parameters of stepwise feature selection were chosen such that too few features were selected by the stepwise procedure. When feature selection was performed using only the design samples, the hold-out estimate was always pessimistically biased. When feature selection was performed using the entire finite sample space, the hold-out estimates could be pessimistically or optimistically biased, depending on the number of features available for selection, the number of available samples, and their statistical distribution. For our simulation conditions, these estimates were always pessimistically (conservatively) biased if the ratio of the total number of available samples per class to the number of available features was greater than five.

Algorithms↗

Metal contents in the groundwater of Sahebgunj district, Jharkhand, India, with special reference to arsenic.

A detailed study has been presented on groundwater metal contents of Sahebgunj district in the state of Jharkhand, India with special reference to arsenic. Both tubewell and well waters have been studied separately with greater emphasis on tubewell waters. Groundwaters of all the nine blocks of Sahebgunj district have been surveyed for iron, manganese, calcium, magnesium, copper and zinc in addition to arsenic. Normal distribution statistic, exploratory data analysis and robust Z-score analysis have been employed to find out the distribution pattern, localisation of data, outliers and other related information. Groundwaters of three blocks of Sahebgunj, namely, Sahebgunj, Rajmahal and Udhawa have been found to be alarmingly contaminated with arsenic present at or above 10 ppb. Arsenic distribution patterns in these blocks are highly asymmetric in nature with the common feature of increasing width from first to fourth quartile. A very broad fourth quartile in each case represents a long asymmetric tail on the right of the median. Tubewell waters of at least two more blocks require regular monitoring to identify the outbreak of arsenic at the onset. Groundwaters of Sahebgunj district in general contain high iron and manganese. It is by and large soft in nature. Well waters have been found to be better with regard to arsenic but iron and manganese contents do not vary significantly. Normal distribution analysis (NDA), box and whisker (BW) plot and Z-score analysis together can provide a reasonably complete statistical picture of metal contents in Sahebgunj district groundwaters.

Arsenic↗