Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Modified nonparametric approaches to detecting differentially expressed genes in replicated microarray experiments.

MOTIVATION: An important goal in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Various parametric tests, such as the two-sample t-test, have been used, but their possibly too strong parametric assumptions or large sample justifications may not hold in practice. As alternatives, a class of three nonparametric statistical methods, including the empirical Bayes method of Efron et al. (2001), the significance analysis of microarray (SAM) method of Tusher et al. (2001) and the mixture model method (MMM) of Pan et al. (2001), have been proposed. All the three methods depend on constructing a test statistic and a so-called null statistic such that the null statistic's distribution can be used to approximate the null distribution of the test statistic. However, relatively little effort has been directed toward assessment of the performance or the underlying assumptions of the methods in constructing such test and null statistics. RESULTS: We point out a problem of a current method to construct the test and null statistics, which may lead to largely inflated Type I errors (i.e. false positives). We also propose two modifications that overcome the problem. In the context of MMM, the improved performance of the modified methods is demonstrated using simulated data. In addition, our numerical results also provide evidence to support the utility and effectiveness of MMM.

Algorithms↗

Assessment of the spatial occurrence of childhood leukaemia mortality using standardized rate ratios with a simple linear Poisson model.

Reports of a suspected cluster of childhood leukaemia cases in West Central Phoenix have led to a number of epidemiological studies in the geographical area. We report here on a death certificate-based mortality study, which indicated an elevated rate ratio of 1.95 during 1966-1986, using the remainder of the Phoenix standard metropolitan statistical area (SMSA) as a comparison region. In the process of analysing the data from this study, a methodology for dealing with denominator variability in a standardized mortality ratio was developed using a simple linear Poisson model. This new approach is seen as being of general use in the analysis of standardized rate ratios (SRR), as well as being particularly appropriate for cluster investigations.

Adolescent↗

Application of pattern recognition techniques to the analysis of protein crystal structure data. I. Characteristics of per-residue side chain contact frequency distributions.

A statistical study of amino acid side chain contact interactions was carried out using a data set based on 36 protein structures. For each type of amino acid, a distribution of per-residue inter-side-chain contacts was obtained, over the observed span of zero to 11 contacts per residue. Significant observations included the following: 1) The mean number of inter-side-chain contacts is proportional to side chain surface area with the exception of Lys and Arg. 2) The mean number of contacts was greater for amino acids in beta-sheet relative to alpha-helical regions. 3) The more polar or surface-loving amino acids exhibited non-normal distributions, whereas distributions for the non-polar or interior-loving amino acids fell within accepted limits of normality.

Amino Acids↗

Measuring effect size: a non-parametric analogue of omega 2.

When comparing two groups of subjects, one of the many measures of effect size is omega 2. This paper suggests a non-parametric analogue of omega 2 based on the following point of view. Given an observation from one of two groups, but not knowing whether it came from the first or second group, how certain can we be that the observation came from the first group? This is in contrast to omega 2 where, given that an observation came from a specific group, say the first group, how much does this reduce our uncertainty about the dependent variable? One problem with omega 2 is that it is not robust--it is a function of the variances--so it can be misleading for reasons reviewed in the paper. Four estimators of the proposed measure of effect size are described and compared in a simulation study. Contrary to what was expected, the .632 bootstrap estimator performed best in terms of bias and mean squared error.

Adult↗

[Births, fertility, rhythms and lunar cycle. A statistical study of 5,927,978 births].

Is there any relationship between the times when babies are born and the synodic lunar cycle? There are published works that show that there is such a relationship. We have looked at 5,927,978 French births occurring between the months of January 1968 and the 31st December 1974. Using Fourier's spectral analysis we have been able to show that there are two different rhythms in birth frequencies: --a weekly rhythm characterised by the lowest number of births on a Sunday and the largest number on a Tuesday: --an annual rhythm with the maximum number of births in May and the minimum in September-October. A statistical analysis of the distribution of births in the lunar month shows that more are born between the last quarter and the new moon, and fewer are born in the first quarter of the moon. The differences between the distribution observed during the lunar month and the theoretical distribution are statistically significant.

Astronomical Phenomena↗

A multiplicative statistical model predicts the size distribution of unruptured intracranial aneurysms.

A statistical model for characterizing the erratic nature of aneurysm evolution is developed and tested. This model is based upon a multiplicative hypothesis, whereby it is theorized that the progressive changes in the size of a given aneurysm are determined by random multipliers. Such a model would predict that within a large population of aneurysms, a lognormal histogram for aneurysm sizes would occur (i.e. the logarithms of aneurysm size would have a normal distribution). When applied to previously published clinical data of unruptured aneurysms by Crompton (1966) and McCormick et al. (1970), the model is found to adequately describe both sets of data. The methods introduced in this paper illustrate the utility of incorporating statistical and clinical insights with fundamental biometry for studying the complex phenomena of aneurysm growth and rupture.

Aneurysm, Ruptured↗

Interaction effects on counting statistics and the transmission distribution.

We investigate the effect of weak interactions on the full counting statistics of charge transfer through an arbitrary mesoscopic conductor. We show that the main effect can be incorporated into an energy dependence of the transmission eigenvalues and study this dependence in a nonperturbative approach. An unexpected result is that all mesoscopic conductors behave at low energies such as either a single or a double tunnel junction, which divides them into two broad classes.

Journal Article↗

Statistical testing and null distributions: what to do when samples are not random.

Selected literature related to statistical testing is reviewed to compare the theoretical models underlying parametric and nonparametric inference. Specifically, we show that these models evaluate different hypotheses, are based on different concepts of probability and resultant null distributions, and support different substantive conclusions. We suggest that cognitive scientists should be aware of both models, thus providing them with a better appreciation of the implications and consequences of their choices among potential methods of analysis. This is especially true when it is recognized that most cognitive science research employs design features that do not justify parametric procedures, but that do support nonparametric methods of analysis, particularly those based on the method of permutation/randomization.

Humans↗

[Analysis of the statistical characteristsics of the distribution of NK/Ly lymphoma cells obtained by laser flow cytometry].

Distributions in volumes of NK/Ly ascite lymphoma cells in different periods after inoculation were studied by laser cytometry. With the age of tumour the primary peak of distribution shifts to an area corresponding to large cells, but from the sixth day there appears an additional peak of distribution due to an accumulation of a pool of nonproliferating quiscent G0(R1) cells. The obtained data are in good agreement with results of sedimentation fractionation of NK/Ly lymphoma cells. Discrepancies in statistical characteristics of distributions, found by the both methods, are not more than 20%.

Animals↗

[Statistical patterns in the distribution of spontaneous subarachnoid hemorrhages].

A study of 788 cases of spontaneous subarachnoidal hemorrhages, which were observed in some neurological clinics during the past 5 years, permitted an analysis of the statistical data concerning the distribution and outcomes of this form of strokes depending upon the age of the patient and etiology. Among the observed population, individuals of a young and middle age prevailed (78% of the patients were younger than 60 years). The role of separate etiological factors is different in different age groups in individuals under 40, subarachnoidal hemorrhages are mainly produced by a rupture of cerebral vascular aneurysms; in the age group of 40--59 years by arterial hypertension, while in the elderly by atherosclerosis in combination with arterial hypertension. The outcomes of subarachnoidal hemorrhages in general are more favourable in the older age groups. However, in hemorrhages of an aneurysmal etiology the lethality amounts to 21%, and is especially high in repeated strokes. In the group of patients with hemorrhages of an aneurysmal etiology the highest lethality was recorded in a localization of the aneurysm in the anterior communicating artery and anterior cerebral artery.

Adult↗

Protein-water displacement distributions.

The statistical properties of fast protein-water motions are analyzed by dynamic neutron scattering experiments. Using isotopic exchange, one probes either protein or water hydrogen displacements. A moment analysis of the scattering function in the time domain yields model-independent information such as time-resolved mean square displacements and the Gauss-deviation. From the moments, one can reconstruct the displacement distribution. Hydration water displays two dynamical components, related to librational motions and anomalous diffusion along the protein surface. Rotational transitions of side chains, in particular of methyl groups, persist in the dehydrated and in the solvent-vitrified protein structure. The interaction with water induces further continuous protein motions on a small scale. Water acts as a plasticizer of displacements, which couple to functional processes such as open-closed transitions and ligand exchange.

Biophysical Phenomena↗

Ties in proximity and clustering compounds.

Hierarchical clustering algorithms such as Wards or complete-link are commonly used in compound selection and diversity analysis. Many such applications utilize binary representations of chemical structures, such as MACCS keys or Daylight fingerprints, and dissimilarity measures, such as the Euclidean or the Soergel measure. However, hierarchical clustering algorithms can generate ambiguous results owing to what is known in the cluster analysis literature as the ties in proximity problem, i.e., compounds or clusters of compounds that are equidistant from a compound or cluster in a given collection. Ambiguous ties can occur when clustering only a few hundred compounds, and the larger the number of compounds to be clustered, the greater the chance for significant ambiguity. Namely, as the number of "ties in proximity" increases relative to the total number of proximities, the possibility of ambiguity also increases. To ensure that there are no ambiguous ties, we show by a probabilistic argument that the number of compounds needs to be less than 2(n 1/4), where n is the total number of proximities, and the measure used to generate the proximities creates a uniform distribution without statistically preferred values. The common measures do not produce uniformly distributed proximities, but rather statistically preferred values that tend to increase the number of ties in proximity. Hence, the number of possible proximities and the distribution of statistically preferred values of a similarity measure, given a bit vector representation of a specific length, are directly related to the number of ties in proximities for a given data set. We explore the ties in proximity problem, using a number of chemical collections with varying degrees of diversity, given several common similarity measures and clustering algorithms. Our results are consistent with our probabilistic argument and show that this problem is significant for relatively small compound sets.

Journal Article↗

Statistical Inference (part II): The Normal and Related Distributions.

The normal (or gaussian) is the most important probability distribution in the statistical inference process. By transforming the normal distribution to a standard distribution (z-distribution) it is possible to determine the probabilities where observations or values fall in certain intervals for variables with different means and standard deviations. Areas under the normal distribution may be represented in a table where the z is the number of standard deviations away from the mean. For example, the area under the normal distribution delimited by a z value of 1.96 to the left and 1.96 to the right of the mean corresponds to 95% of the total area of the normal distribution. Sampling distribution of estimates of population parameters may also be described by the normal distribution. It is important to note, however, that the estimation of a population mean based on the normal distribution is conditional to the assumption that the population standard deviation is a known parameter. The t distribution is used to infer about a population mean when the population standard deviation is estimated by the sample data. In statistical inference about proportions, the normal approximation of the binomial distribution may be used provided the data fit certain assumptions. Statistical methods that do not depend on the form of the distribution (distribution-free or nonparametric methods) and those based on the actual probability distribution, called exact methods, are often used in situations where the data do not comply with the assumption that the distribution of the estimate is approximately normal.

Journal Article↗