Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Testing for trend with count data.

Among the tests that can be used to detect dose-related trends in count data from toxicological studies are nonparametric tests such as the Jonckheere-Terpstra and likelihood-based tests, for example, based on a Poisson model. This paper was motivated by a data set of tumor counts in which conflicting conclusions were obtained using these two tests. To define situations where one test may be preferable, we compared the small and large sample performance of these two tests as well as a robust and conditional version of the likelihood-based test in the absence and presence of a dose-related trend for both Poisson and overdispersed Poisson data. Based on our results, we suggest using the Poisson test when little overdispersion is present in the data. For more overdispersed data, we recommend using the robust Poisson test for highly discrete data (response rate lower than 2-3) and the robust Poisson test or the Jonckheere-Terpstra test for moderately discrete or continuous data (average responses larger than 2 or 3). We also studied the effects of dose metameter misspecification. A clear effect on efficiency was seen when the 'wrong' dose metameter was used to compute the test statistic. In general, unless there is strong reason to do otherwise, we recommend the use of equally spaced dose levels when applying the Poisson or robust Poisson test for trend.

Animals↗

Statistical analysis of the surface distribution of microtubule-associated proteins (MAPs) bound in vitro to rat brain mitochondria and labelled by 10 nm gold-coupled antibodies.

Purified mitochondria from rat brain were incubated in vitro which microtubule-associated proteins (MAPs) that are known to bind specifically on sites present on the outer membrane. The bound molecules were detected by immunoelectron microscopy and the linear distribution of the label along mitochondrial profiles was analyzed by statistical methods. The results demonstrate that gold-conjugated antibodies are distributed in a non-random fashion on the surface of mitochondria, suggesting regional concentrations of MAPs-binding sites. This finding argue for the existence of specialized domains on mitochondria that are involved in the association of the organelles to microtubules in situ.

Animals↗

Factors affecting the statistical parameters and patterns of distribution of residual moistures in arrays of samples following lyophilization.

Numerous studies were undertaken to examine in depth the statistical parameters (mean, standard deviation, standard error of the mean, etc.) and the distributions of residual moistures of 3 mL samples of 2% albumin arranged in a 22 x 10 grid of samples (the array) and dried by sublimation of ice in vacuo. Because of the relatively large numbers of selected samples used for the determination of the patterns of distribution of residual moistures, it was necessary to develop a modified gravimetric method for measuring residual moistures. In the majority of circumstances, with the vials containing the samples positioned on a base defined by cartesian coordinates (x, y) and the contents of residual moistures measured on an axis perpendicular to the base (z), the distributions of contents of residual waters were described best by the equation for a 2-nd order polynomial. It was found that the contents of residual moistures of individual vials were "position oriented" and not "vial associated." It was also found that the statistical parameters and patterns of distributions of residual moistures were dependent on: (1) the orientation of openings in the split stoppers used for freeze-drying, (2) the shelf of the freeze-drying apparatus (top, middle, bottom) upon which the vials were placed, and (3) alterations in elapsed times and shelf temperatures.

Animals↗

The large sample distribution of the weighted log rank statistic under general local alternatives.

We derive the large sample distribution of the weighted log rank statistic under a general class of local alternatives in which both the cure rates and the conditional distribution of time to failure among those who fail are assumed to vary in the two treatment arms. The analytic result presented here is important to data analysts who are designing clinical trials for diseases such as non-Hodgkins lymphoma, leukemia and melanoma, where a significant proportion of patients are cured. We present a numerical illustration comparing powers obtained from the analytic result to those obtained from simulations.

Clinical Trials as Topic↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗

On statistical tests of phylogenetic tree imbalance: the Sackin and other indices revisited.

We investigate the distribution of statistical measures of tree imbalance in large phylogenies. More specifically, we study normalized versions of the Sackin's index and the number of subtrees of given sizes. Using the connection with structures from theoretical computer science, we provide precise description for the limiting distribution under the null hypothesis of Yule trees. Corrected p-values are then computed, and the statistical power of these statistics for testing the Yule model against a model of biased speciation is evaluated from simulations. As an illustration, the tests are applied to the HIV-1 reconstructed phylogeny.

Acquired Immunodeficiency Syndrome↗

On the chi-square approximation to the exact distribution of goodness-of-fit statistics in multinomial models with composite hypotheses.

Multinomial models are increasingly being used in psychology, and this use always requires estimating model parameters and testing goodness of fit with a composite null hypothesis. Goodness of fit is customarily tested with recourse to the asymptotic approximation to the distribution of the statistics. An assessment of the quality of this approximation requires a comparison with the exact distribution, but how to compute this exact distribution when parameters are estimated from the data appears never to have been defined precisely. The main goal of this paper is to compare two different approaches to defining this exact distribution. One of the approaches uses the marginal distribution and is, therefore, independent of the data; the other approach uses the conditional distribution of the statistics given the estimated parameters and, therefore, is data-dependent. We carried out a thorough study involving various parameter estimation methods and goodness-of-fit statistics, all of them members of the general class of power-divergence measures. Included in the study were multinomial models with three to five cells and up to three parameters. Our results indicate that the asymptotic distribution is rarely a good approximation to the exact marginal distribution of the statistics, whereas it is a good approximation to the exact conditional distribution only when the vector of expected frequencies is interior to the sample space of the multinomial distribution.

Binomial Distribution↗

Cholesky problems.

Behavioral geneticists commonly parameterize a genetic or environmental covariance matrix as the product of a lower diagonal matrix postmultiplied by its transpose-a technique commonly referred to as "fitting a Cholesky." Here, simulations demonstrate that this procedure is sometimes valid, but at other times: (1) may not produce fit statistics that are distributed as a chi2; or (2) if the distribution of the fit statistic is chi2, then the degrees of freedom (df) are not always the difference between the number of parameters in the general model less the number of parameters in a constrained model. It is hypothesized that the problem is related to the fact that the Cholesky parameterization requires that the covariance matrix formed by the product be either positive definite or singular. Even though a population covariance matrix may be positive definite, the combination of sampling error and the derived--as opposed to directly observed--nature of genetic and environmental matrices allow matrices that are negative (semi) definite. When this occurs, fitting a Cholesky constrains the numerical area of search and compromises the maximum likelihood theory currently used in behavioral genetics. Until the reasons for this phenomenon are understood and satisfactory solutions are developed, those who fit Cholesky matrices face the burden of demonstrating the validity of their fit statistics and the df for model comparisons. An interim remedy is proposed--fit an unconstrained model and a Cholesky model, and if the two differ, then report the difference in fit statistics and parameter estimates. Cholesky problems are a matter of degree, not of kind. Thus, some Cholesky solutions will differ trivially from the unconstrained solutions, and the importance of the problems must be assessed by how often the two lead to different substantive interpretation of the results. If followed, the proposed interim remedy will develop a body of empirical data to assess the extent to which Cholesky problems are important substantive issues versus statistical curiosities.

Algorithms↗

Statistical analysis of corticopontine neuron distribution in visual areas 17, 18, and 19 of the cat.

The spatial organization of visual corticopontine neurons was studied both at a "large scale" (in relation to cortical visual field maps) and at a "small scale" (in relation to cortical modular organization). Large injections of horse-radish peroxidase-wheat germ agglutinin were made in the pontine nuclei. In complete series of sections from parts of areas 17, 18, and 19, the position of each retrogradely labeled neuron was recorded with an x-y plotter connected to the microscope stage. Each cell was thus given a set of x, y, and z coordinates. After alignment of the sections, three-dimensional computer reconstructions of the distribution of the labeled cells were made. With program RPOP (developed by Blackstad and Bjaalie, '88), the reconstructions were studied with different rotations, scaling, etc. In addition, section-independent parts of reconstructions were isolated ("windows") and further analyzed. Curved parts were automatically unfolded for inspection of distribution patterns and determination of cell densities. The spatial distribution of the labeled cells was analyzed within small windows, where density gradients are negligible. We confirm and extend previous demonstrations of a large-scale aggregation of visual corticopontine cells due to density gradients by showing that densities of corticopontine neurons increase linearly as a function of distance from paracentral to lower visual field representations in area 17 (and partly in areas 18 and 19). We demonstrate that density gradients are steeper in area 17 than in area 18. For example, clear-cut differences between the areas in mediolateral density gradients are found. These findings are discussed in relation to the different visual field maps of the areas and the existence of a similar visual field representation in corticopontine projections from different visual areas. The type of small-scale distribution (randomness or non-randomness, aggregation into clusters, bands, etc.) was studied with statistical methods. Such analysis shows that the labeled cells within small zones are non-randomly distributed in all three areas. In most cases, the analysis indicates an aggregated spatial distribution. A possible relationship to the cortical map of direction selectivity is discussed. To our knowledge, this study is the first to combine the use of three-dimensional computer reconstructions of a population of labeled neurons, with subsequent statistical analysis of spatial point (cell distribution) patterns.

Animals↗

A test for linkage and association in general pedigrees: the pedigree disequilibrium test.

Family-based tests of linkage disequilibrium typically are based on nuclear-family data including affected individuals and their parents or their unaffected siblings. A limitation of such tests is that they generally are not valid tests of association when data from related nuclear families from larger pedigrees are used. Standard methods require selection of a single nuclear family from any extended pedigrees when testing for linkage disequilibrium. Often data are available for larger pedigrees, and it would be desirable to have a valid test of linkage disequilibrium that can use all potentially informative data. In this study, we present the pedigree disequilibrium test (PDT) for analysis of linkage disequilibrium in general pedigrees. The PDT can use data from related nuclear families from extended pedigrees and is valid even when there is population substructure. Using computer simulations, we demonstrated validity of the test when the asymptotic distribution is used to assess the significance, and examined statistical power. Power simulations demonstrate that, when extended pedigree data are available, substantial gains in power can be attained by use of the PDT rather than existing methods that use only a subset of the data. Furthermore, the PDT remains more powerful even when there is misclassification of unaffected individuals. Our simulations suggest that there may be advantages to using the PDT even if the data consist of independent families without extended family information. Thus, the PDT provides a general test of linkage disequilibrium that can be widely applied to different data structures.

Alleles↗

Variability of birth-weight distributions by sex and ethnicity: analysis using mixture models.

Birth weight is the most important proximate determinant of the level of infant mortality. However, the association between birth weight and infant mortality is not constant among populations. For example, the mortality of African American infants is lower at low birth weight but higher at high birth weight compared with European American infants. One possible explanation is that birth cohorts are heterogeneous even after controlling for birth weight, ethnicity, sex, and multiple births. The analyses presented here use Gaussian mixture models to explore the interpopulation variation in the shape of the birth-weight distribution for evidence of intrapopulation heterogeneity. The results suggest that a two-component mixture model provides an excellent description of human birth-weight distributions. Further statistical analyses of sex and ethnic differences indicate (1) that the birth-weight distributions and heterogeneity within the distribution vary between the sexes and among ethnic groups and (2) that one specific component is more closely associated with the overall level of infant mortality. The results support the hypothesis that birth cohorts can consist of two or more subpopulations at differential risk of mortality. Differences in the subpopulation composition of birth cohorts (i.e., differences in the level of heterogeneity among the various ethnic groups) might partially explain the interethnic variation in birth-weight-specific mortality. Further development of these mixture models should provide important additional information concerning the biological, environmental, and social determinants of birth weight and infant mortality.

Black or African American↗

Description of atomic burials in compact globular proteins by Fermi-Dirac probability distributions.

We perform a statistical analysis of atomic distributions as a function of the distance R from the molecular geometrical center in a nonredundant set of compact globular proteins. The number of atoms increases quadratically for small R, indicating a constant average density inside the core, reaches a maximum at a size-dependent distance R(max), and falls rapidly for larger R. The empirical curves turn out to be consistent with the volume increase of spherical concentric solid shells and a Fermi-Dirac distribution in which the distance R plays the role of an effective atomic energy epsilon(R) = R. The effective chemical potential mu governing the distribution increases with the number of residues, reflecting the size of the protein globule, while the temperature parameter beta decreases. Interestingly, betamu is not as strongly dependent on protein size and appears to be tuned to maintain approximately half of the atoms in the high density interior and the other half in the exterior region of rapidly decreasing density. A normalized size-independent distribution was obtained for the atomic probability as a function of the reduced distance, r = R/R(g), where R(g) is the radius of gyration. The global normalized Fermi distribution, F(r), can be reasonably decomposed in Fermi-like subdistributions for different atomic types tau, F(tau)(r), with Sigma(tau)F(tau)(r) = F(r), which depend on two additional parameters mu(tau) and h(tau). The chemical potential mu(tau) affects a scaling prefactor and depends on the overall frequency of the corresponding atomic type, while the maximum position of the subdistribution is determined by h(tau), which appears in a type-dependent atomic effective energy, epsilon(tau)(r) = h(tau)r, and is strongly correlated to available hydrophobicity scales. Better adjustments are obtained when the effective energy is not assumed to be necessarily linear, or epsilon(tau)*(r) = h(tau)*r(alpha,), in which case a correlation with hydrophobicity scales is found for the product alpha(tau)h(tau)*. These results indicate that compact globular proteins are consistent with a thermodynamic system governed by hydrophobic-like energy functions, with reduced distances from the geometrical center, reflecting atomic burials, and provide a conceptual framework for the eventual prediction from sequence of a few parameters from which whole atomic probability distributions and potentials of mean force can be reconstructed.

Amino Acids↗

Statistics for studying quanta at synapses: resampling and confidence limits on histograms.

This paper describes some statistical methods for working with data on quantal sizes. Since quantal sizes often do not fit to normal probability distribution functions, statistics based on the normal distribution are inappropriate. Resampling methods can be used to determine confidence limits and to test whether two sets of data differ by chance. Some hypotheses about the nature of quanta are based on apparent peaks and valleys in histograms. Confidence limits can be placed on the bins in the histogram by using the Kolomorogov-Smirnov statistic or by resampling. The confidence limits should assist in the evaluation of the significance of the valleys.

Animals↗

Modeling the occurrence of cardiac arrest as a poisson process.

STUDY OBJECTIVE: A statistical model for the occurrence of cardiac arrest has not been described in the literature. Independent events occurring along the time axis may constitute a Poisson process, described by the Poisson and exponential probability distributions. This statistical model defines the probability distribution of events occurring within time intervals and enables construction of confidence intervals for the mean rate. Moreover, the probability that 2 or more events will occur close in time can be estimated from knowledge of the mean rate. We investigated whether the occurrence of cardiac arrests constitutes a Poisson process. METHODS: Time and date for cardiac arrests requiring CPR out of hospital (county population, 155,000) or in hospital (850 beds) during 5 years were analyzed. Goodness of fit was assessed by comparing the observed weekly counts of cardiac arrests and the observed time intervals between such events with the values predicted from the model. RESULTS: The Poisson parameter estimates (mean weekly rates) for out-of-hospital and in-hospital cardiac arrest were 2.02 and 1.09 events per week, respectively. There was close agreement between observed and predicted values, indicating an adequate model fit. CONCLUSION: Occurrence of cardiac arrest along the time axis constitutes a Poisson process and may be adequately modeled by the Poisson and exponential distributions. The model provides information about the nature of these events and allows for probability calculations based on the mean rate of events. Examples of such calculations are given.

Cardiopulmonary Resuscitation↗