Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Distributional regimes for the number of k-word matches between two random sequences.

When comparing two sequences, a natural approach is to count the number of k-letter words the two sequences have in common. No positional information is used in the count, but it has the virtue that the comparison time is linear with sequence length. For this reason this statistic D(2) and certain transformations of D(2) are used for EST sequence database searches. In this paper we begin the rigorous study of the statistical distribution of D(2). Using an independence model of DNA sequences, we derive limiting distributions by means of the Stein and Chen-Stein methods and identify three asymptotic regimes, including compound Poisson and normal. The compound Poisson distribution arises when the word size k is large and word matches are rare. The normal distribution arises when the word size is small and matches are common. Explicit expressions for what is meant by large and small word sizes are given in the paper. However, when word size is small and the letters are uniformly distributed, the anticipated limiting normal distribution does not always occur. In this situation the uniform distribution provides the exception to other letter distributions. Therefore a naive, one distribution fits all, approach to D(2) statistics could easily create serious errors in estimating significance.

Computer Simulation↗

Statistical methods for short-term projections of AIDS incidence.

Short-term projections of AIDS incidence are critical for assessing future health care needs. This paper focuses on the method of back-calculation for obtaining short-term projections. The approach consists of back-calculating from AIDS incidence data through use of the incubation period distribution to obtain estimates of the numbers previously infected. The numbers previously infected are then projected forward to obtain short-term projections. An approach is suggested for accounting for new infections in short-term projections of AIDS incidence. Back-calculation requires accurate AIDS incidence data. A method which is computationally easy to implement is proposed for estimating the distribution of the delays in reporting AIDS cases. It was found that the reporting delay distribution in the United States varies by geographic region of diagnosis. Back-calculation also requires a reliable estimate of the incubation period distribution. Statistical issues associated with estimating the incubation period distribution are considered. The methods are applied to obtain short-term projections of AIDS incidence in the United States. The projected cumulative AIDS incidence in the U.S. by the end of 1992 was 287,100 under the assumption that there are no new infections after 1 July 1987, and 330,600 under the assumption that the infection rate remains constant. These projections do not account for the new broadened AIDS surveillance definitions or the underreporting of AIDS cases to the Centers for Disease Control.

Acquired Immunodeficiency Syndrome↗

A statistical thermodynamic model applied to experimental AFM population and location data is able to quantify DNA-histone binding strength and internucleosomal interaction differences between acetylated and unacetylated nucleosomal arrays.

Imaging of nucleosomal arrays by atomic force microscopy allows a determination of the exact statistical distributions for the numbers of nucleosomes per array and the locations of nucleosomes on the arrays. This precision makes such data an excellent reference for testing models of nucleosome occupation on multisite DNA templates. The approach presented here uses a simple statistical thermodynamic model to calculate theoretical population and positional distributions and compares them to experimental distributions previously determined for 5S rDNA nucleosomal arrays (208-12,172-12). The model considers the possible locations of nucleosomes on the template, and takes as principal parameters an average free energy of interaction between histone octamers and DNA, and an average wrapping length of DNA around the octamers. Analysis of positional statistics shows that it is possible to consider interactions between nucleosomes and positioning effects as perturbations on a random positioning noninteracting model. Analysis of the population statistics is used to determine histone-DNA association constants and to test for differences in the free energies of nucleosome formation with different types of histone octamers, namely acetylated or unacetylated, and different DNA templates, namely 172-12 or 208-12 5S rDNA multisite templates. The results show that the two template DNAs bind histones with similar affinities but histone acetylation weakens the association of histones with both templates. Analysis of locational statistics is used to determine the strength of specific nucleosome positioning tendencies by the DNA templates, and the strength of the interactions between neighboring nucleosomes. The results show only weak positioning tendencies and that unacetylated nucleosomes interact much more strongly with one another than acetylated nucleosomes; in fact acetylation appears to induce a small anticooperative occupation effect between neighboring nucleosomes.

Acetylation↗

Quantal measurement and analysis methods compared for crayfish and Drosophila neuromuscular junctions, and rat hippocampus.

Quantal content of transmission was estimated for three synaptic systems (crayfish and Drosophila neuromuscular junctions, and rat dentate gyrus neurons) with three different methods of measurement: direct counts of released quanta, amplitude measurements of evoked and spontaneous events, and charge measurements of evoked and spontaneous events. At the crayfish neuromuscular junction, comparison of the three methods showed that estimates from charge measurements were closer to estimates from direct counts, since amplitude measurements were more seriously affected by variable latency in evoked release of quantal units. Thus, charge measurements are better for estimating quantal content when direct counts cannot be made, as in crayfish at high frequency of stimulation or in the dentate gyrus neurons. At the Drosophila neuromuscular junction, there is almost no latency variation of quantal release in realistic physiological solutions, and the methods based upon amplitudes and charge give similar results. Distributions of evoked synaptic quantal events obtained by direct counts at the crayfish neuromuscular junction were compared to statistical distributions obtained by best fits. Binomial distributions with uniform or non-uniform probabilities of release generally provided good fits to the observations. From best fit distributions, the quantal parameters n (number of release sites) and p (their probability of release) can be calculated. We used two algorithms to estimate n and p: one allows for non-uniform probability of release and uses a modified chi-square (chi 2) criterion, and the second assumes uniform probability of release and derives parameters from maximum likelihood estimation (MLE). The bootstrap estimate of standard errors is used to determine the accuracy of n and p estimates.

Animals↗

Epigenetic randomness, complexity and singularity of human iris patterns.

We investigated the randomness and uniqueness of human iris patterns by mathematically comparing 2.3 million different pairs of eye images. The phase structure of each iris pattern was extracted by demodulation with quadrature wavelets spanning several scales of analysis. The resulting distribution of phase sequence variation among different eyes was precisely binomial, revealing 244 independent degrees of freedom. This amount of statistical variability corresponds to an entropy (information density) of about 3.2 bits mm(-2) over the iris. It implies that the probability of two different irides agreeing by chance in more than 70% of their phase sequence is about one in 7 billion. We also compared images of genetically identical irides, from the left and right eyes of 324 persons, and from monozygotic twins. Their relative phase sequence variation generated the same statistical distribution as did unrelated eyes. This indicates that apart from overall form and colour, iris patterns are determined epigenetically by random events in the morphogenesis of this tissue. The resulting diversity, and the combinatorial complexity created by so many dimensions of random variation, mean that the failure of a simple test of statistical independence performed on iris patterns can serve as a reliable rapid basis for automatic personal identification.

Functional Laterality↗

Oxygen fields in specific spinal loci of the canine spinal cord.

Oxygen tension (PO2) measurements were made in the dog spinal cord with a small recessed-tip oxygen microelectrode. The use of vibration and specific marking techniques has allowed the elimination of tissue compression artifacts and the mapping of regional PO2 in the thoracic spinal cord. A symmetrical distribution of PO2 values can be shown for the lateral white funiculi; the gray matter and dorsal columns have multimodal distributions. Statistical evaluation showed all these areas to have different PO2 profiles; PO2 values (mmHg) were 61.2 +/- 12.4 for the lateral white funiculi, 55.3 +/- 19.0 for the dorsal columns, and 30.0 +/- 13.6 in spinal gray. The relatively normal distribution patterns of these oxygen tensions indicate that traditional statistical methods may be used to compare and evaluate oxygen diffusion fields in the adult spinal cord.

Animals↗

Single-particle tracking: the distribution of diffusion coefficients.

In single-particle tracking experiments, the diffusion coefficient D may be measured from the trajectory of an individual particle in the cell membrane. The statistical distribution of single-trajectory diffusion coefficients is examined by Monte Carlo calculations. The width of this distribution may be useful as a measure of the heterogeneity of the membrane and as a test of models of hindered diffusion in the membrane. For some models, the distribution of the short-range diffusion coefficient is much narrower than the observed distribution for proteins diffusing in cell membranes. To aid in the analysis of single-particle tracking measurements, the distribution of D is examined for various definitions of D and for various trajectory lengths.

Animals↗

Guidelines for the statistical evaluation of SCE.

When planning studies by the sister chromatid exchange (SCE) test, it is necessary to calculate the size of the test groups, taking into consideration the variance of the test result and the statistical distribution of SCE frequencies. This paper deals with these problems. Recommendations are given for the preparation of slides under standardized conditions, for the subsequent counting of SCEs/cell in a random sample of 30 cells from each slide, and for the condensation of the information contained in the sample of 30 counts into a single statistic, that may be treated as a normally distributed variable. The adequacy of this transformation is shown for the data from 170 different subjects. Of these, 165 (58 nonsmokers and 107 cigarette smokers) had mean values of SCE/cell ranging from 6.5 to 13.5, while the remaining 5 subjects were on intermittent treatment with cytostatics every fourth wk, and exhibited a mean value of SCE/cell in the range 15-23. The variance associated with the recommended statistic has been decomposed into 4 variance components: variance within slides, variance between slides prepared from the same blood sample, variance within subjects, and variance between subjects. Based on a total of 680 SCE analyses in 218 persons, estimates of these variance components are given and used to calculate the necessary number of samples for the detection of a prescribed difference in SCEs/cell for selected values of Type I and Type II errors.

Analysis of Variance↗

Modelling the size separated particulate matter (SSPM10) from vehicular exhaust at traffic intersections in Mumbai.

The study was carried out to predict the size separated particulate matter below 10 microm size (SSPM10) from vehicular exhausts at traffic intersections using modified general finite line source model (GFLSM). Two air quality control regions (AQCRs) were selected in Mumbai City for this study. One was industrial area (AQCR1) containing the busy intersection, i.e. Marol link road, with the heavy inflow of two-three wheelers. And, the other was commercial busy district area (AQCR2) containing the busy intersection, i.e. Dadar circle, with a heavy traffic flow especially cars. The model was applied at both the traffic intersections. The data were collected for modelling study for three winter months in 1995 using cascade impactor of nine size ranges. The prediction results revealed that modified GFLSM underpredicted the SSPM10 concentrations for all the size ranges. However, showed considerable correlation between observed and predicted values for the size range below 4.7 microm at both the intersections. The relative high concentrations observed in the coarser range of 10-4.7 microm are attributed to the resuspension of the roadside particulate matter. Hence, the amount of underprediction was more for this range, which was due to the characteristics of model that does not take into account the factor for resuspension of roadside particulate matter caused by traffic movements. The model was also applied to predict the total particulate matter for downwind distances from the road intersection. The statistical evaluation of model was done, which indicated that the model's performance was good for the finer range of particles (below 4.7 microm) with r-square values of 0.49 and 0.57 found at both the intersections in AQCR1 and AQCR2, respectively. However, it is not unusual that the model uncertainty is likely to exist due to data input errors and stochastic fluctuations irrespective of the models accurateness. The statistical distribution model was therefore identified using Kolmogorov-Smirnov test. At both the intersections, SSPM10 concentration data were found lognormally distributed.

Air Pollutants↗

Distribution of individual cytoplasmic pH values in a population of the yeast Saccharomyces cerevisiae.

Fluorescence ratio imaging microscopy using pH-sensitive fluorescent dyes makes it possible to evaluate statistical distribution of intracellular pH in a population of the yeast S. cerevisiae examined in a thin layer of suspension in a Petri dish. The distribution appears to fit a Gaussian curve with a half-width around the 0.4 pH unit. The curve became slightly narrower after resuspension in a strong buffer; the mean values shifted with the pH of the buffer. The shape of the distribution curves of both resting and growing cells in various phases of growth does not change significantly. Likewise, addition of 1% of glucose, 50 microM suloctidil or 100 microM diethylstilbestrol brings about no alteration. The only value which clearly changes is the average cytoplasmic pH.

Cytoplasm↗

An analysis of alternative classification schemes for medical atlas mapping.

Most national disease atlases adopt a classification scheme based on either the percentile distribution of rates or on the national mean. Although these schemes have a direct interpretation, they are based on the univariate statistical distribution of rates and not on their spatial distribution, and distort the underlying spatial autocorrelation in the data. If the purpose of the maps is to represent spatial patterns, alternative classification schemes might be more appropriate. This research proposes an alternative classification method that maximises spatial similarity among contiguous units in the same class interval. The method has been illustrated using selected data from the German Cancer Atlas published in 1984.

Colonic Neoplasms↗

Distribution-based criteria for change in health-related quality of life in Parkinson's disease.

BACKGROUND AND OBJECTIVE: To be useful, results from health-related quality of life (HRQoL) measures must be interpretable. The objective of this article is to examine statistical (distributional) approaches to interpretability. The standard error of measurement (SEM) and the standard error of the difference (S(diff)) are used in data on individuals with Parkinson's disease to calculate the minimum change scores required to be statistically meaningful for each dimension of an instrument to assess HRQoL in Parkinson's disease, the PDQ-39. METHODS: Data was collected from both a community and a clinic study; in both studies the PDQ-39 was administered at baseline and follow-up. RESULTS: The patterns of SEMs and S(diff)s were similar both across time periods and between samples, for all dimensions except Social Support. CONCLUSIONS: The results suggest that, for example, six points change on a 0-100 transformed scoring of the Mobility dimension may be considered on distributional grounds a minimum meaningful change. The demonstrated consistency across occasions and types of sample of SEMs and S(diff) for the majority of the dimensions of the PDQ-39, is evidence of the theoretically claimed advantage of this measure of sample independence, and supports use of this distributional approach to minimum meaningful change.

Adult↗

On cortical folds and neuromagnetic fields.

A folded cortical source of neuromagnetic fields, similar in configuration to the visual cortex, was simulated. Cortical activity was modelled by different distributions of independent current dipoles. The map of the summed fields of the dipoles of this cruciform model changed, depending upon the statistical distribution of the electrical activity of the dipoles and its geometry. Arrays of dipoles of random orientations and strengths produced field patterns that could be interpreted as due to moving neural currents, although the geometry of the neural tissue remained unchanged and the average activity remained approximately constant. The field topography at any instant was apparently unrelated to the depth or orientation of the underlying structure, thus raising questions about how to interpret topographic MEG and EEG displays. Furthermore, asynchronous activity (defined as independent directions and magnitudes of activity of the dipoles) did not result in less field power than when the dipoles were synchronized, i.e., when the direction of current flow was correlated across all of the dipoles within the cruciform structure. Therefore, in this model 'alpha blockage' cannot be mimicked by desynchronization. More generally, for the cruciform or any other symmetrically folded and active cortical sheet, 'blockage' cannot be attributed to desynchronization. The same is true for the EEG except that smooth unfolded sheets of radially oriented dipoles would result in enhancement of voltage due to synchronization. Such radial dipoles do not contribute to the MEG. Blockage was simulated by reducing the amount of activity within different portions of the synchronized cruciform model. This resulted in a dramatic increase in the net field because attenuation broke the symmetry of the synchronized cruciform structure. With asynchronous dipoles populating the structure, the attenuation of the same portion of the structure had no easily discerned effect on the net field. However, maps of average field power were consistently related to the position of the region of attenuated activity. The locations of regions of attenuated activity were determined by taking the difference between the mean square field pattern obtained when all portions of the cruciform structure were active and the pattern obtained when a portion of the structure was relatively inactive. When activity of the same portions were incremented rather than attenuated, the resulting plot of average power was essentially the same as that of the attenuated portion derived by taking these differences between power distributions. The major conclusions are that the concepts of synchronization and desynchronization have no explanatory power unless the physical conditions under which they occur are specified precisely.(ABSTRACT TRUNCATED AT 400 WORDS)

Brain↗

On the statistical significance of nucleic acid similarities.

When evaluating sequence similarities among nucleic acids by the usual methods, statistical significance is often found when the biological significance of the similarity is dubious. We demonstrate that the known statistical properties of nucleic acid sequences strongly affect the statistical distribution of similarity values when calculated by standard procedures. We propose a series of models which account for some of these known statistical properties. The utility of the method is demonstrated in evaluating high relative similarity scores in four specific cases in which there is little biological context by which to judge the similarities. In two of the cases we identify the statistical properties which are responsible for the apparent similarity. In the other two cases the statistical significance of the similarity persists even when the known statistical properties of sequences are modelled. For one of these cases biological significance is likely while the other case remains an enigma.

Base Sequence↗

Power vectors: an application of Fourier analysis to the description and statistical analysis of refractive error.

The description of sphero-cylinder lenses is approached from the viewpoint of Fourier analysis of the power profile. It is shown that the familiar sine-squared law leads naturally to a Fourier series representation with exactly three Fourier coefficients, representing the natural parameters of a thin lens. The constant term corresponds to the mean spherical equivalent (MSE) power, whereas the amplitude and phase of the harmonic correspond to the power and axis of a Jackson cross-cylinder (JCC) lens, respectively. Expressing the Fourier series in rectangular form leads to the representation of an arbitrary sphero-cylinder lens as the sum of a spherical lens and two cross-cylinders, one at axis 0 degree and the other at axis 45 degrees. The power of these three component lenses may be interpreted as (x,y,z) coordinates of a vector representation of the power profile. Advantages of this power vector representation of a sphero-cylinder lens for numerical and graphical analysis of optometric data are described for problems involving lens combinations, comparison of different lenses, and the statistical distribution of refractive errors.

Computer Graphics↗

Analysis of gene duplication repeats in the myosin rod.

The helical coiled-coil region of the myosin rod in the nematode Caenorhabditis elegans is a repetitive sequence 1094 amino acids long which contains 39 repeats of a 28-residue pattern. The repeats are extremely significant when compared with the statistical distributions expected, first for random sequences, and then for sequences with a typical seven-residue coiled-coil periodicity. New and improved statistical tests are used. The repeats are stronger in the first 350 residues of the rod (fragment S-2) than in the remainder. The corresponding DNA sequence of the unc-54 gene shows the same features, but they are less significant when judged by the number of identical bases than are the amino acid similarities, as measured by Dayhoff scores. The rod sequence shows strong evidence for a longer repeat unit of 196 residues, which may be related to the cross-bridge spacing of 143 A in muscle.

Amino Acid Sequence↗

The relationship between personality and attainment in 16-19-year-old students in a sixth form college. I: Construction of the Student Self-Perception Scale.

BACKGROUND: Of the research that has been undertaken into the relationship between personality and attainment, relatively little exists relating to the 16-19 age range. In a substantive study examining the relationship between academic self-concept, attainment and personality in sixth form students, a first requirement was to design a self-perception instrument. AIMS: The psychometric element of the study aimed to construct a Student Self-Perception Scale (SSPS) that would be effective for students in the FE (further education) context. SAMPLES: The samples comprised a pilot sample of 152 students (aged 16-17 years from two sixth from colleges) and a main sample of 364 students (mean age, 16yrs 10mths, range 16:0 to 18:6 years, from one sixth form college). The main sample included similar numbers of male and female students (46% male, 54% female) and ethnic minority students comprised 14% of this sample. METHOD: An initial item pool of 88 four-point Likert type statements was compiled from comparable existing scales and from responses to a Student Induction Questionnaire. Item analysis was based on oblique factor analysis of the pilot sample responses, followed by cross-validation on the main sample to refine the scale structures. Construct validity was established from the substantive study, especially the Nowicki & Strickland (1973) locus of control results. RESULTS: Exploration of the four- and five-factor structures led to a final specification based on 52 items from five oblique factors. The constituent scales were Passivity (12 items, alpha = .81). Mastery (15 items, alpha = .79), Work Related Inadequacy (11 items, alpha = .72), Extraversion (4 items, alpha = .70) and Social Dependence (10 items, alpha = .66), all statistics compiled from the cross-validation sample. Correlations with Locus of Control ranged from 0.52 for Mastery to -.34 for Work Related Inadequacy. Distribution statistics for Locus of Control matched a comparable American sample. CONCLUSIONS: The five-scale structure exhibits good cross-validation characteristics and supports revealing analyses of relationships within the substantive study. Its 52-item format is suitable for research or exploratory use within its intended FE context.

Achievement↗