Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Construction of null statistics in permutation-based multiple testing for multi-factorial microarray experiments.

MOTIVATION: The parametric F-test has been widely used in the analysis of factorial microarray experiments to assess treatment effects. However, the normality assumption is often untenable for microarray experiments with small replications. Therefore, permutation-based methods are called for help to assess the statistical significance. The distribution of the F-statistics across all the genes on the array can be regarded as a mixture distribution with a proportion of statistics generated from the null distribution of no differential gene expression whereas the other proportion of statistics generated from the alternative distribution of genes differentially expressed. This results in the fact that the permutation distribution of the F-statistics may not approximate well to the true null distribution of the F-statistics. Therefore, the construction of a proper null statistic to better approximate the null distribution of F-statistic is of great importance to the permutation-based multiple testing in microarray data analysis. RESULTS: In this paper, we extend the ideas of constructing null statistics based on pairwise differences to neglect the treatment effects from the two-sample comparison problem to the multifactorial balanced or unbalanced microarray experiments. A null statistic based on a subpartition method is proposed and its distribution is employed to approximate the null distribution of the F-statistic. The proposed null statistic is able to accommodate unbalance in the design and is also corrected for the undue correlation between its numerator and denominator. In the simulation studies and real biological data analysis, the number of true positives and the false discovery rate (FDR) of the proposed null statistic are compared with those of the permutated version of the F-statistic. It has been shown that our proposed method has a better control of the FDRs and a higher power than the standard permutation method to detect differentially expressed genes because of the better approximated tail probabilities.

Algorithms↗

[Hydrolytic enzyme activities in the bottom sediment cores from Norwegian Sea and statistical analysis of their distribution].

Proteinase and amylase enzyme activities were evaluated in bottom sediment cores from the Norwegian Sea collected along a transect from the summit plane of the Voring Plateau on the east to fault uplifts of the Yan-Mayen transform zone perpendicular to the present-day Norwegian Current. Spotted vertical distribution of hydrolytic enzyme activities by the location and depth of the cores and specific distribution of proteinase and amylase activities have been revealed in four bottom sediment cores (up to 300 cm; 5 cm resolution). Specific activity distribution has been revealed for different types of enzyme-sorbing bottom sediments. Current methods of statistical analysis and mathematical modeling were applied to reveal the relationship between enzymatic degradation of protein and polysaccharide organic compounds and the content of carbonates and organic matter in bottom sediments.

Amylases↗

[The suitability of the graphic decomposition methods of Daeves and Beckel for mixed distributions in nuclear variation statistics].

The object of these studies was to determine the value of the graphic method of DAEVES and BECKEL (1958) for decomposition of mixed distributions in karyometry and to compare this method with a numerical one, published by HEROLD (1971). The former is as good as the latter, the only disadvantage of the graphic method is the loss of time by reason of the fact that employment of computers is not possible.

Aged↗

Maximally selected chi-square statistics for ordinal variables.

The association between a binary variable Y and a variable X having an at least ordinal measurement scale might be examined by selecting a cutpoint in the range of X and then performing an association test for the obtained 2 x 2 contingency table using the chi-square statistic. The distribution of the maximally selected chi-square statistic (i.e. the maximal chi-square statistic over all possible cutpoints) under the null-hypothesis of no association between X and Y is different from the known chi-square distribution. In the last decades, this topic has been extensively studied for continuous X variables, but not for non-continuous variables of at least ordinal measurement scale (which include e.g. classical ordinal or discretized continuous variables). In this paper, we suggest an exact method to determine the finite-sample distribution of maximally selected chi-square statistics in this context. This novel approach can be seen as a method to measure the association between a binary variable and variables having an at least ordinal scale of different types (ordinal, discretized continuous, etc). As an illustration, this method is applied to a new data set describing pregnancy and birth for 811 babies.

Biometry↗

A statistical estimator of the spatial distribution of the water-table altitude.

An algorithm was designed to statistically estimate the areal distribution of water-table altitude. The altitude of the water table was bounded below by the minimum water-table surface and above by the land surface. Using lake elevations and stream stages, and interpolating between lakes and streams, the minimum water-table surface was generated. A multiple linear regression among the minimum water-table altitude, the differerence between land-surface and minimum water-table altitudes, and the water-level measurements from surficial aquifier system wells resulted in a consistently high correlation for all groups of physiographic regions in Florida. A simple linear regression between land-surface and water-level measurements resulted in a root-mean-square residual of 4.23 m, with residuals ranging from -8.78 to 41.54 m. A simple linear regression between the minimum water table and the water-level measurements resulted in a root-mean-square residual of 1.45 m, with residuals ranging from -7.39 to 4.10 m. The application of the multiple linear regression presented herein resulted in a root-mean-square residual of 1.05 m, with residuals ranging from -5.24 to 5.63 m. Results from complete and partial F tests rejected the hypothesis of eliminating any of the regressors in the multiple linear regression presented in this study.

Algorithms↗

An affine invariant rank-based method for comparing dependent groups.

A basic property of various rank-based hypothesis testing methods is that they are invariant under a linear transformation of the data. For multivariate data, a generalization of this property is sometimes sought (called affine invariance), but typically techniques for assigning ranks do not achieve this goal, or it is assumed that sampling is from a symmetric distribution. A rank-based method is suggested for comparing dependent groups that is based on halfspace depth, is affine invariant in terms of difference scores, and allows sampling from asymmetric distributions.

Analysis of Variance↗

[Investigations on the individual-region distribution of adipocyte diameters by means of advanced statistical methods].

The dimensional distributions of the adipocytes in Equus caballus in many subjects and in many regions have been studied: such distributions turn out to be in good approximation galtonian ones. Furthermore, all the logarithm populations of the cell diameters have significantly the same variance. The used statistical methods (ANOVA two way with replications, and TUKEY -test) indicate an extremely significant different among the various regions (the smallest cells are in the supra-orbital fossa, the greatest ones are in the abdominal subserous floor).

Adipose Tissue↗

Three-parameter lognormal distribution ubiquitously found in cDNA microarray data and its application to parametric data treatment.

BACKGROUND: To cancel experimental variations, microarray data must be normalized prior to analysis. Where an appropriate model for statistical data distribution is available, a parametric method can normalize a group of data sets that have common distributions. Although such models have been proposed for microarray data, they have not always fit the distribution of real data and thus have been inappropriate for normalization. Consequently, microarray data in most cases have been normalized with non-parametric methods that adjust data in a pair-wise manner. However, data analysis and the integration of resultant knowledge among experiments have been difficult, since such normalization concepts lack a universal standard. RESULTS: A three-parameter lognormal distribution model was tested on over 300 sets of microarray data. The model treats the hybridization background, which is difficult to identify from images of hybridization, as one of the parameters. A rigorous coincidence of the model to data sets was found, proving the model's appropriateness for microarray data. In fact, a closer fitting to Northern analysis was obtained. The model showed inconsistency only at very strong or weak data intensities. Measurement of z-scores as well as calculated ratios was reproducible only among data in the model-consistent intensity range; also, the ratios were independent of signal intensity at the corresponding range. CONCLUSION: The model could provide a universal standard for data, simplifying data analysis and knowledge integration. It was deduced that the ranges of inconsistency were caused by experimental errors or additive noise in the data; therefore, excluding the data corresponding to those marginal ranges will prevent misleading analytical conclusions.

Blotting, Northern↗

Lattice models, packing density, and Boltzmann-like distribution of cavities in proteins.

A model reproducing the experimental Boltzmann-like distribution of empty cavity sizes in proteins is introduced. Proteins are represented by lattices of different dimensionalities, corresponding to different numbers of nearest neighbor contacts. Small cavities emerge and join into larger ones in a random process that can be related to random mutations. Simulations of cavity creation are performed under the constraint of a limiting total packing density. Cavities sufficiently large (20 A(3) or more), that they might accommodate at least one additional methyl group produced by a mutation, are counted and compared to the distribution of cavities according to their sizes from protein statistics. The distributions calculated with this very simple model within a realistic range of packing densities are in good agreement with the empirical cavity distribution. The results suggest that the Boltzmann-like distribution of cavities in proteins might be affected by a mechanism controlled by limiting packing density and maximum allowed protein destabilization. This supports an earlier suggestion that the agreement between the free energies of cavity formation from the mutational experiments and from the statistics of the empty cavity distribution in X-ray protein structures is nonfortuitous. A possible relation of the suggested model to the Boltzmann hypothesis is discussed.

Computational Biology↗

Sudden and unexpected deaths after the administration of hexavalent vaccines (diphtheria, tetanus, pertussis, poliomyelitis, hepatitis B, Haemophilius influenzae type b): is there a signal?

UNLABELLED: Deaths in temporal association with vaccination of hexavalent vaccines have been recently reported. The objective of this paper is to assess whether these temporal associations can be attributed to chance. Standardised mortality ratios (SMR) for deaths within 1 to 28 days after administration of either of the two hexavalent vaccines in the 1st and 2nd year of life were determined using the respective annual rates for sudden unexpected deaths (SUDs) from the national vital statistics. The distribution of SUD cases and the vaccination uptake by month were estimated from surveys and sales figures for the individual vaccines. Sensitivity analyses were performed to account for limitations in the data sources. For one of the vaccines, Vaccine B, all SMRs were well below one. For the other, Vaccine A, SMRs exceeded one insignificantly on the 1st day after vaccination in the 1st year of life. In the 2nd year of life, however, the SMRs for SUD cases within 1 day of vaccination with vaccine A were 31.3 (95% CI 3.8-113.1; two cases observed; 0.06 cases expected) and 23.5 (95% CI 4.8-68,6) for within 2 days after vaccination (three cases observed; 0.13 cases expected). Extensive sensitivity analyses could not attribute these findings to limitations of the data sources. CONCLUSION: These findings based on spontaneous reporting do not prove a causal relationship between vaccination and sudden unexpected deaths. However, they constitute a signal for one of the two hexavalent vaccines which should prompt intensified surveillance for unexpected deaths after vaccination.

Age Distribution↗

Direct measurement of single and ensemble average particle-surface potential energy profiles.

This work involves the development of a novel technique that integrates total internal reflection and video microscopy methods to simultaneously measure single particle and ensemble average particle-surface interactions. For the 2 mum silica colloids and glass coverslip used in this study, particle size polydispersity is found to be a dominant factor in determining the distribution of single particle profiles about ensemble average profiles. In conjunction with this observation, chemical and physical nonuniformity are not evident in any of our measurements even with sensitivity to interactions on the order of kT. One advantage of using ensemble averaging in conjunction with time averaging is the ability to dramatically decrease the time required to measure average particle-wall interactions which scales inversely with interfacial particle concentration. A number of experimental issues are addressed in the development of this technique including (1) combining single particle distribution functions, (2) statistical sampling of distribution functions using both time and ensemble averaging, and (3) correcting overlapping scattering signals between adjacent particles. The capabilities of the ensemble averaging technique are also demonstrated to provide unique measurements of particle-surface interactions in metastable systems by selecting only height excursions of levitated particles when calculating potentials. Ultimately, this new technique provides several important advantages over single particle measurements, which provides a foundation for measuring interactions in increasingly complex interfacial systems.

Journal Article↗

Molecular thermodynamics for swelling of a mesoscopic ionomer gel in 1 : 1 salt solutions.

For a microphase-separated diblock copolymer ionic gel swollen in salt solution, a molecular-thermodynamic model is based on the self-consistent field theory in the limit of strongly segregated copolymer subchains. The geometry of microdomains is described using the Milner generic wedge construction neglecting the packing frustration. A geometry-dependent generalized analytical solution for the linearized Poisson-Boltzmann equation is obtained. This generalized solution not only reduces to those known previously for planar, cylindrical and spherical geometries, but is also applicable to saddle-like structures. Thermodynamic functions are expressed analytically for gels of lamellar, bicontinuous, cylindrical and spherical morphologies. Molecules are characterized by chain composition, length, rigidity, degree of ionization, and by effective polymer-polymer and polymer-solvent interaction parameters. The model predicts equilibrium solvent uptakes and the equilibrium microdomain spacing for gels swollen in salt solutions. Results are given for details of the gel structure: distribution of mobile ions and polymer segments, and the electric potential across microdomains. Apart from effects obtained by coupling the classical Flory-Rehner theory with Donnan equilibria, viz. increased swelling with polyelectrolyte charge and shrinking of gel upon addition of salt, the model predicts the effects of microphase morphology on swelling.

Algorithms↗

Common noncompartmental pharmacokinetic variables: are they normally or log-normally distributed?

We investigated the hypothesis that distributions of continuous pharmacokinetic variables are positively skewed in nature and that logarithmic transformation of these variables restores normality. The distributions of common continuous noncompartmental pharmacokinetic variables were investigated for four different Glaxo Wellcome compounds, administered by three different routes of administration: ranitidine (po), sumatriptan (sc), ondansetron (iv), and bismuth, from ranitidine bismuth citrate (po). The distributions of all the investigated noncompartmental pharmacokinetic variables were adequately described by a log-normal distribution, whereas statistically significant departures from normality occurred in the majority of cases. Thus, unless there is strong and consistent evidence for a departure from log-normality, the parametric statistical analysis of common noncompartmental pharmacokinetic variables should be carried out after a priori log transformation.

Bismuth↗

An entropy-based statistic for genomewide association studies.

Efficient genotyping methods and the availability of a large collection of single-nucleotide polymorphisms provide valuable tools for genetic studies of human disease. The standard chi2 statistic for case-control studies, which uses a linear function of allele frequencies, has limited power when the number of marker loci is large. We introduce a novel test statistic for genetic association studies that uses Shannon entropy and a nonlinear function of allele frequencies to amplify the differences in allele and haplotype frequencies to maintain statistical power with large numbers of marker loci. We investigate the relationship between the entropy-based test statistic and the standard chi2 statistic and show that, in most cases, the power of the entropy-based statistic is greater than that of the standard chi2 statistic. The distribution of the entropy-based statistic and the type I error rates are validated using simulation studies. Finally, we apply the new entropy-based test statistic to two real data sets, one for the COMT gene and schizophrenia and one for the MMP-2 gene and esophageal carcinoma, to evaluate the performance of the new method for genetic association studies. The results show that the entropy-based statistic obtained smaller P values than did the standard chi2 statistic.

Entropy↗

The empirical association between student and resident physician performances.

To further the understanding of the relationship between performances in a combined baccalaureate-MD degree program and in residency, the authors subjected their database of 298 study participants from the 1980-1983 entering classes of the University of Missouri-Kansas City School of Medicine to factor analysis and then to distribution-free statistical analyses. Distinct factors were identified among the performance measures from the combined-degree program; only one factor was identified among the measures of residency performance. Analysis of the relationship of performances in the combined-degree program and in residency indicated that almost half of the participants were in the same performance categories as students and as residents. The strongest association emerged between a clinical performance factor derived from performances in the combined-degree program and the residency clinical performance factor. However, the knowledge factor derived from performance measures in the combined-degree program was also associated with residency clinical performance. The associations were statistically significant but of limited strength; thus the present results resemble those of other investigators, despite the fact that they are based on distribution-free statistics and on purportedly cleaner and more homogeneous measures of performance. Various technical reasons may have caused this lack of strength, but it is also possible that empirical relationships between undergraduate and postgraduate performances are inherently limited because the performances expected of residents may not be mere extensions of those expected of medical students.

Achievement↗