Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Phenylketonuria mutations and linked haplotypes in the Lithuanian population: origin of the most common R408W mutation.

A genealogical study was performed in Lithuanian phenylketonuria (PKU) families with the aim of tracing the origins of the R408W/haplotype 2/VNTR3 allele. The relative frequency of six phenylalanine hydroxylase (PAH) mutations (R408W, R158Q, R261Q, G272X, IVS10nt-11g --> a, and IVS12nt1g --> a) common in Eastern European populations and their association with variable number of tandem repeat (VNTR) and short tandem repeat (STR) sites in the PAH gene were examined in 130 PKU Lithuanian chromosomes, including 95 of Baltic, 28 of Slavonic and 7 of unknown origin. R408W was found to be the most frequent (70%) mutation in both Balts or Slavonians with a uniform frequency distribution. No statistically significant differences in the frequency distribution of the other mutations analysed were found. In Balts and Slavonians, the R408W mutation is strongly associated with the three-copy VNTR and the 240-bp STR allele. The frequency of this association is 68% in both ethnic groups. The genealogical data provided in this paper indicate that the most common R408W/VNTR3/STR240 allele arose in ancient times possibly among pre-Indo-Europeans and suggest that the high frequency of the R408W mutation and associated minihaplotype in Balts of Lithuania is due to a founder effect.

Evolution, Molecular↗

An iterative method for improved estimation of the mean of peer-group distributions in proficiency testing.

In proficiency testing (PT), the peer-group mean is conventionally computed after twice removing values exceeding the mean +/- 3 SD. However, this adjustment fails if there are many outliers. In this study an iterative method was evaluated as a more robust way to estimate the means. The methodology repeatedly removes a proportion of the population (usually those exceeding the mean +/- 1.6 SD), assuming the presence of a Gaussian distribution in the central portion, and reinflates the SD to compensate for the trimming. A computer simulation revealed that the estimated mean of a known Gaussian distribution was less affected by a subpopulation that overlaps the main population than was the conventional method. When the overlapping portions were removed, the iterative method predicted the true mean correctly. The method was applied to external PT results for 44 analytes. Although most peer-group distributions were clearly non-Gaussian, the segment included by the predicted mean +/- 1.6 SD was regarded as Gaussian in 85.9% by the new method and 73.4% by the conventional method. The proposed methodology appears to be an improved way of estimating peer-group means.

Algorithms↗

Analysis of overdispersed count data by mixtures of Poisson variables and Poisson processes.

Count data often show overdispersion compared to the Poisson distribution. Overdispersion is typically modeled by a random effect for the mean, based on the gamma distribution, leading to the negative binomial distribution for the count. This paper considers a larger family of mixture distributions, including the inverse Gaussian mixture distribution. It is demonstrated that it gives a significantly better fit for a data set on the frequency of epileptic seizures. The same approach can be used to generate counting processes from Poisson processes, where the rate or the time is random. A random rate corresponds to variation between patients, whereas a random time corresponds to variation within patients.

Anticonvulsants↗

Remote sensing and spatial statistical analysis to predict the distribution of Oncomelania hupensis in the marshlands of China.

Remote sensing and spatial statistical analysis were employed to predict the distribution of Oncomelania hupensis, the intermediate host snail of Schistosoma japonicum, in the marshlands of Jiangning county in China. Surrogate indices related to environmental factors in the marshlands were derived from a Landsat 7 ETM+ image, and the relationship between environmental covariates and the density of O. hupensis was analyzed by stepwise regression models and ordinary kriging. Although stepwise regression demonstrated that O. hupensis densities of live snails in the marshlands related significantly to the modified soil-adjusted vegetation index, wetness and land surface temperature, the correlation coefficient was low (0.282). Therefore, spatial patterns of the regression residual were investigated by the semi-variogram method, and the spatial variation of O. hupensis density attributed to the spatial autocorrelation was estimated by ordinary kriging. The regression model of the snail density and ordinary kriging of its spatial variation were then combined with the aim of improving the prediction of O. hupensis. Following this approach, the prediction indeed improved considerably (0.852). Our results show that it is possible to predict the distribution of O. hupensis in these marshlands by using remotely sensed environmental indices, and that spatial statistical analyses are capable of improving prediction accuracy. These findings are of relevance for mapping and prediction of schistosomiasis japonica in China, and hence the national control programme.

Animals↗

[Rank distributions in community ecology from the statistical viewpoint].

Traditional statistical methods for definition of empirical functions of abundance distribution (population, biomass, production, etc.) of species in a community are applicable for processing of multivariate data contained in the above quantitative indices of the communities. In particular, evaluation of moments of distribution suffices for convolution of the data contained in a list of species and their abundance. At the same time, the species should be ranked in the list in ascending rather than descending population and the distribution models should be analyzed on the basis of the data on abundant species only.

Animals↗

On the distribution of the unpaired t-statistic with paired data.

I derive the exact distribution of the unpaired t-statistic computed when the data actually come from a paired design. I use this to prove a result Diehr et al. obtained by simulation, namely that the type I error rate of this procedure is no greater than alpha regardless of the sample size. I provide a formula to use in computation of power and type I error rate.

Analysis of Variance↗

A novel locus for clubroot resistance in Brassica rapa and its linkage markers.

An inbred turnip ( Brassica rapa syn. campestris) line, N-WMR-3, which carries the trait of clubroot resistance (CR) from a European turnip, Milan White, was crossed with a clubroot-susceptible doubled haploid line, A9709. A segregating F(3) population was obtained by single-seed descent of F(2) plants and used for a genetic analysis. Segregation of CR in the F(3) population suggested that CR is controlled by a major gene. Two RAPD markers, OPC11-1 and OPC11-2, were obtained as candidates of linkage markers by bulked segregant analysis. These were converted to sequence-tagged site markers, by cloning and sequencing of the polymorphic bands, and named OPC11-1S and OPC11-2S, respectively. The specific primer pairs for OPC11-1S amplified a clear dominant band, while the primer pairs for OPC11-2S resulted in co-dominant bands. Frequency distributions and statistical analyses indicate the presence of a major dominant CR gene linked to these two markers. The present marker for CR was independent of the previously found CR loci, Crr1 and Crr2. Genotypic distribution and statistical analyses did not show any evidence of CR alleles on Crr1 and Crr2 loci in N-WMR-3. The present study clearly demonstrates that B. rapa has at least three CR loci. Therefore, the new CR locus was named Crr3. The present locus may be useful in breeding CR Chinese cabbage cultivars to overcome the decay of present CR cultivars.

Brassica rapa↗

Comparisons of continuous and discrete methods for combining probability values associated with matched-pairs t-test data.

Fisher's well-known continuous method for combining independent probability values from continuous distributions is compared with an exact discrete analog of Fisher's continuous method for combining independent probability values from discrete distributions using matched-pairs t-test data. Fisher's continuous method is shown to be inadequate for combining probability values from many discrete distributions, given the continuity assumption when discrete distributions are considered. Although Fisher's continuous method does not detect a well-documented effect among distributions, the exact discrete analog method clearly detects the effect.

Algorithms↗

Multiple testing. Part II. Step-down procedures for control of the family-wise error rate.

The present article proposes two step-down multiple testing procedures for asymptotic control of the family-wise error rate (FWER): the first procedure is based on maxima of test statistics (step-down maxT), while the second relies on minima of unadjusted p-values (step-down minP). A key feature of our approach is the characterization and construction of a test statistics null distribution (rather than data generating null distribution) for deriving cut-offs for these test statistics (i.e., rejection regions) and the resulting adjusted p-values. For general null hypotheses, corresponding to submodels for the data generating distribution, we identify an asymptotic domination condition for a null distribution under which the step-down maxT and minP procedures asymptotically control the Type I error rate, for arbitrary data generating distributions, without the need for conditions such as subset pivotality. Inspired by this general characterization, we then propose as an explicit null distribution the asymptotic distribution of the vector of null value shifted and scaled test statistics. Step-down procedures based on consistent estimators of the null distribution are shown to also provide asymptotic control of the Type I error rate. A general bootstrap algorithm is supplied to conveniently obtain consistent estimators of the null distribution.

Journal Article↗

Gaussian distributions on Lie groups and their application to statistical shape analysis.

The Gaussian distribution is the basis for many methods used in the statistical analysis of shape. One such method is principal component analysis, which has proven to be a powerful technique for describing the geometric variability of a population of objects. The Gaussian framework is well understood when the data being studied are elements of a Euclidean vector space. This is the case for geometric objects that are described by landmarks or dense collections of boundary points. We have been using medial representations, or m-reps, for modelling the geometry of anatomical objects. The medial parameters are not elements of a Euclidean space, and thus standard PCA is not applicable. In our previous work we have shown that the m-rep model parameters are instead elements of a Lie group. In this paper we develop the notion of a Gaussian distribution on this Lie group. We then derive the maximum likelihood estimates of the mean and the covariance of this distribution. Analogous to principal component analysis of covariance in Euclidean spaces, we define principal geodesic analysis on Lie groups for the study of anatomical variability in medially-defined objects. Results of applying this framework on a population of hippocampi in a schizophrenia study are presented.

Algorithms↗

[Relationship between the fine structure and statistic fluctuations in measurement result distributions].

The "near zone effect" described first in the macroscopic fluctuations researches was studied by numerical modeling. Possible mechanisms of its formation were proposed, and the lengths of portions of experimental data that retain the fine structure of histograms necessary for studying the "near zone effect" and provide the minimum level of statistical noise.

Models, Theoretical↗

Monte Carlo dose calculations and radiobiological modelling: analysis of the effect of the statistical noise of the dose distribution on the probability of tumour control.

The aim of this work is to investigate the influence of the statistical fluctuations of Monte Carlo (MC) dose distributions on the dose volume histograms (DVHs) and radiobiological models, in particular the Poisson model for tumour control probability (tcp). The MC matrix is characterized by a mean dose in each scoring voxel, d, and a statistical error on the mean dose, sigma(d); whilst the quantities d and sigma(d) depend on many statistical and physical parameters, here we consider only their dependence on the phantom voxel size and the number of histories from the radiation source. Dose distributions from high-energy photon beams have been analysed. It has been found that the DVH broadens when increasing the statistical noise of the dose distribution, and the tcp calculation systematically underestimates the real tumour control value, defined here as the value of tumour control when the statistical error of the dose distribution tends to zero. When increasing the number of energy deposition events, either by increasing the voxel dimensions or increasing the number of histories from the source, the DVH broadening decreases and tcp converges to the 'correct' value. It is shown that the underestimation of the tcp due to the noise in the dose distribution depends on the degree of heterogeneity of the radiobiological parameters over the population; in particular this error decreases with increasing the biological heterogeneity, whereas it becomes significant in the hypothesis of a radiosensitivity assay for single patients, or for subgroups of patients. It has been found, for example, that when the voxel dimension is changed from a cube with sides of 0.5 cm to a cube with sides of 0.25 cm (with a fixed number of histories of 10(8) from the source), the systematic error in the tcp calculation is about 75% in the homogeneous hypothesis, and it decreases to a minimum value of about 15% in a case of high radiobiological heterogeneity. The possibility of using the error on the tcp to decide how many histories to run for a given voxel size is also discussed.

Computer Simulation↗