Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Temporal distribution of death among oncology patients: environmental links.

BACKGROUND: The natural history of terminal oncologic disease is death from cardiopulmonary arrest. The goal of this study was to determine whether the temporal distribution of death among oncology patients is related to levels of environmental physical activity-solar, geomagnetic, and cosmic rays and high energy space proton flux, as previously shown for other situations like accidents, suicide, and the occurrence of acute myocardial infarction. EXPERIMENTAL: Deaths of oncology patients, n = 102604, in 168 consecutive months were compared with-monthly indices of solar, geomagnetic, cosmic rays activity (sunspot number, solar flux, Ap, Cp, Am), and indices of magnetic and cosmic rays activity according to neutron monitor data). In addition, oncology patients monthly death numbers were compared with numbers of deaths from ischemic heart disease (IHD), stroke, sudden cardiac death (SCD), accidents, road accidents, non cardiovascular death (total deaths - [IHD + stroke]). The National Data of the Republic of Lithuania was used, excluding SCD and myocardial infarction in Kaunas--the second largest city in Lithuania Monica study register was used for SCD at age 25-64, and all data of AMI. Pearson correlation coefficients and their probabilities were obtained between the numbers of oncology patient monthly deaths and (1) cosmo-physical indices and (2) the number of deaths from other causes and the occurrence of AMI. RESULTS: The number of oncology patient deaths was inversely correlated with solar and geomagnetic activity indices and positively correlated with cosmic rays activity. The number of oncology patient deaths was correlated with monthly number of deaths from non-cardiovascular courses, stroke, suicide, at trend level with SCD, but not with deaths from IHD, accidents, road accidents. Oncology patient deaths showed a significant correlation with the number of AMI. CONCLUSION: The monthly death number of oncology patients is significantly related with environmental physical activity and shows similarity with deaths distribution by time of some other groups of death--stroke, SCD, suicide, and occurrence of AMI. The inverse correlation with solar and geomagnetic activity and positive links with cosmic rays activity level are remarkable.

Accidents↗

[A comparison of complete and truncated birth histories to measure fertility and child mortality].

"During the latter part of 1986, national probability sample surveys of women of reproductive ages were carried out in... Peru and the Dominican Republic. These surveys were made as part of the Demographic Health Surveys project (DHS). In each country, one survey was conducted with the standard core questionnaire developed for DHS; the other survey was based on an experimental questionnaire. The major difference between the two questionnaires is the inclusion in the experimental one of a monthly calendar, which records pregnancies, contraceptive use, reasons for contraceptive discontinuation, breastfeeding, post-partum amenorrhea, post-partum abstinence, women's employment and place of residence for the period 1981-1986. This paper presents results from the first stage of the analysis of the Peruvian data: a comparison of basic characteristics of the two samples and an assessment of the completeness of reporting of recent births and infant and child deaths, i.e., a comparison of information in the truncated and full birth histories." (SUMMARY IN ENG)

Americas↗

Shape modelling using Markov random field restoration of point correspondences.

A method for building statistical point distribution models is proposed. The novelty in this paper is the adaption of Markov random field regularization of the correspondence field over the set of shapes. The new approach leads to a generative model that produces highly homogeneous polygonized shapes and improves the capability of reconstruction of the training data. Furthermore, the method leads to an overall reduction in the total variance of the point distribution model. Thus, it finds correspondence between semi-landmarks that are highly correlated in the shape tangent space. The method is demonstrated on a set of human ear canals extracted from 3D-laser scans.

Algorithms↗

Forms and distribution of selenium at different depths and among particle size fractions of three Taiwan soils.

The bioavailability of selenium in soils for plants depends more on its forms than on its total concentration. The purpose of the present study was to examine the solid-phase forms of selenium at different depths of three soil series representing major farming soil groups in Taiwan as well as the amounts of selenium in sand, silt and clay fractions of the soils. The study was conducted by means of sequential extraction to obtain the amounts of selenium and the distribution of various solid-phase forms of selenium at different depths of Pinchen (121 degrees 11(')E, 24 degrees 55(')N), Toulun-Sheto (120 degrees 55(')E, 24 degrees 50(')N), and Chunliao (120 degrees 25(')E, 23 degrees 57(')N) soil series. The amounts of metal oxide-bound form of selenium in the three soil series were the largest, with those of Pinchen and Toulun-Sheto soil series exceeding 50% of the total amounts of selenium and that of Chunliao soil series maintained at 30-40%. In the Pinchen and Toulun-Sheto soil series, the amounts of selenium in clay fractions were the largest, with a significant difference between the clays with and without metal oxides and organic matter removed. The amounts of selenium remained high in silt and/or sand fractions of the Chunliao soil series with metal oxides and organic matter removed. Metal oxide and organic matter contents of the three soil series mainly affect the amounts of various solid-phase forms of selenium and their distribution in different depths and particle size fractions of the soils. This observation of selenium associated with soil constituents was in good agreement with the results of the adsorption of selenite and selenate by the three soil series.

Adsorption↗

Analysis of the frequency distribution of tuberculin skin test readings: a tool for the assessment of group contact investigations.

SETTING: The public health tuberculosis control program covering Seattle, Washington, and its surrounding suburban areas. OBJECTIVE: To describe a tool of potential usefulness in the assessment of transmission of tuberculosis in contact investigations of groups, such as co-workers or schoolmates of an infectious case, with a low prior probability of latent tuberculosis infection (LTBI). DESIGN: Tuberculin skin test (TST) readings in mm of the group being tested were graphed and compared with the known frequency distributions of TST readings of populations with and without LTBI, the latter including a fraction with non-specific tuberculin reactivity. RESULTS: Four group contact investigations were analyzed retrospectively with this tool. In two the graphed TST readings of contacts fell within the distribution of a population with LTBI, and suggested that transmission had occurred. In the other two, the graphed readings better fit the distribution of a population with non-specific tuberculin reactivity and suggested that transmission had not occurred. CONCLUSION: This simple tool to facilitate the determination of whether transmission of tuberculosis has occurred, and who should be offered treatment for LTBI in contact investigations of groups of people, deserves further study.

Adolescent↗

Particle size distributions of organic aerosol constituents during the 2002 Yosemite Aerosol Characterization Study.

The Yosemite Aerosol Characterization Study (YACS) was conducted in the summer of 2002 to investigate sources of regional haze in Yosemite National Park. Organic carbon and molecular source marker species size distributions were investigated during hazy and clear periods. More than 75% of the organic carbon mass was associated with submicron aerosol particles. Most molecular marker species for wood smoke, an important source of particulate matter during the study, were contained in submicron particles, although on some fire influenced days, levoglucosan shifted toward larger sizes. Various wood smoke marker species exhibited slightly different size distributions in the samples, suggesting different, size dependent emission or atmospheric processing rates of these species. Secondary biogenic compounds including pinic and pinonic acids were associated with smaller particles. Pinonaldehyde, however, exhibited a broader distribution, likely due to its higher volatility. Dicarboxylic acids were associated mainly with submicron particles. Hopanes, molecular markers for vehicle emissions, were mostly contained in smaller particles but exhibited some tailing into larger size classes.

Acids↗

Genetic and nongenetic bases for the L-shaped distribution of quantitative trait loci effects.

The L-shaped distribution of estimated QTL effects (R(2)) has long been reported. We recently showed that a metabolic mechanism could account for this phenomenon. But other nonexclusive genetic or nongenetic causes may contribute to generate such a distribution. Using analysis and simulations of an additive genetic model, we show that linkage disequilibrium between QTL, low heritability, and small population size may also be involved, regardless of the gene effect distribution. In addition, a comparison of the additive and metabolic genetic models revealed that estimates of the QTL effects for traits proportional to metabolic flux are far less robust than for additive traits. However, in both models the highest R(2)'s repeatedly correspond to the same set of QTL.

Linkage Disequilibrium↗

On the choice of a sparse prior.

An emerging paradigm analyses in what respect the properties of the nervous system reflect properties of natural scenes. It is hypothesized that neurons form sparse representations of natural stimuli: each neuron should respond strongly to some stimuli while being inactive upon presentation of most others. For a given network, sparse representations need fewest spikes, and thus the nervous system can consume the least energy. To obtain optimally sparse responses the receptive fields of simulated neurons are optimized. Algorithmically this is identical to searching for basis functions that allow coding for the stimuli with sparse coefficients. The problem is identical to maximizing the log likelihood of a generative model with prior knowledge of natural images. It is found that the resulting simulated neurons share most properties of simple cells found in primary visual cortex. Thus, forming optimally sparse representations is a very compact approach to describing simple cell properties. Many ways of defining sparse responses exist and it is widely believed that the particular choice of the sparse prior of the generative model does not significantly influence the estimated basis functions. Here we examine this assumption more closely. We include the constraint of unit variance of neuronal activity, used in most studies, into the objective functions. We then analyze learning on a database of natural (cat-cam) visual stimuli. We show that the effective objective functions are largely dominated by the constraint, and are therefore very similar. The resulting receptive fields show some similarities but also qualitative differences. Even for coefficient values for which the objective functions are dissimilar, the distributions of coefficients are similar and do not match the priors of the assumed generative model. In conclusion, the specific choice of the sparse prior is relevant, as is the choice of additional constraints, such as normalization of variance.

Algorithms↗

Graphical interpretation of confidence curves in rankit plots.

A well-known transformation from the bell-shaped Gaussian (normal) curve to a straight line in the rankit plot is investigated, and a tool for evaluation of the distribution of reference groups is presented. It is based on the confidence intervals for percentiles of the calculated Gaussian distribution and the percentage of cumulative points exceeding these limits. The process is to rank the reference values and plot the cumulative frequency points in a rankit plot with a logarithmic (In=log(e)) transformed abscissa. If the distribution is close to In-Gaussian the cumulative frequency points will fit to the straight line describing the calculated In-Gaussian distribution. The quality of the fit is evaluated by adding confidence intervals (CI) to each point on the line and calculating the percentage of points outside the hyperbola-like CI-curves. The assumption was that the 95% confidence curves for percentiles would show 5% of points outside these limits. However, computer simulations disclosed that approximate 10% of the series would have 5% or more points outside the limits. This is a conservative validation, which is more demanding than the Kolmogorov-Smirnov test. The graphical presentation, however, makes it easy to disclose deviations from In-Gaussianity, and to make other interpretations of the distributions, e.g., comparison to non-Gaussian distributions in the same plot, where the cumulative frequency percentage can be read from the ordinate. A long list of examples of In-Gaussian distributions of subgroups of reference values from healthy individuals is presented. In addition, distributions of values from well-defined diseased individuals may show up as In-Gaussian. It is evident from the examples that the rankit transformation and simple graphical evaluation for non-Gaussianity is a useful tool for the description of sub-groups.

Blood Chemical Analysis↗

Should we maintain the 95 percent reference intervals in the era of wellness testing? A concept paper.

The reference interval is probably the most widely used decision-making tool in clinical practice, with a modern use aiming at identifying wellness during health check and screening. Its use as a diagnostic tool is much less recognised and may be obsolete. The present study investigates the consequences of the new practice for the interpretation of prospective value, negative vs. positive, the probability of confirming wellness, and number of false results based on selected strategy for reference interval establishment. Calculations assumed normalised Gaussian-distributed reference intervals with analytical variation set to zero and absolute accuracy. Also assumed is the independency of tests. Probability for no values outside reference intervals in healthy subjects was calculated from the formula p(no) outside=(1 - p(single)) and according to the formula for repeated testing: p(one) outside =n x p(single) (1 - p(single))n-1 etc. Here n is the number of tests performed and p(single) is the probability of one result outside reference limits with the general formula p(i) outside n-i=k x p(single)i (1- p(single))n-i, with k being the binominal coefficient and i the number outside the reference intervals. Use of the 99.9 centile for health checks will increase the probability for no false from 60% to 99% for 10 tests, and from 46% to 98% for 15 tests. The probability for one false-positive result in 10 tests in a panel can be reduced from 32% to 1% if the 99.9% centile is substituted for the 95% centile. For two in 10 tests, the probability can be reduced from 8% to below 0.1%. In both cases, selection of the 99.9% centile improves the diagnostic accuracy. Reference intervals are needed as a "true" negative reference for absence of disease, and should cover the 99.9% centile of the reference distribution of an analyte to avoid false positives. For this new use, it is critical that reference persons are absolutely normal without clinical, genetic and biochemical signs of the condition being investigated. However, reference intervals cannot substitute clinical decision limits for diagnosis and medical intervention.

Clinical Laboratory Techniques↗

Ethnicity and glutathione S-transferase (GSTM1/GSTT1) polymorphisms in a Brazilian population.

The distribution of polymorphisms related to glutathione S-transferases (GST) has been described in different populations, mainly for white individuals. We evaluated the distribution of GST mu (GSTM1) and theta (GSTT1) genotypes in 594 individuals, by multiplex PCR-based methods, using amplification of the exon 7 of CYP1A1 gene as an internal control. In São Paulo, 233 whites, 87 mulattos, and 137 blacks, all healthy blood-donor volunteers, were tested. In Bahia, where black and mulatto populations are more numerous, 137 subjects were evaluated. The frequency of the GSTM1 null genotype was significantly higher among whites (55.4%) than among mulattos (41.4%; P = 0.03) and blacks (32.8%; P < 0.0001) from São Paulo, or Bahian subjects in general (35.7%; P = 0.0003). There was no statistically different distribution among any non-white groups. The distribution of GSTT1 null genotype among groups did not differ significantly. The agreement between self-reported and interviewer classification of skin color in the Bahian group was low. The interviewer classification indicated a gradient of distribution of the GSTM1 null genotype from whites (55.6%) to light mulattos (40.4%), dark mulattos (32.0%) and blacks (28.6%). However, any information about race or ethnicity should be considered with caution regarding the bias introduced by different data collection techniques, specially in countries where racial admixture is intense, and ethnic definition boundaries are loose. Because homozygous deletions of GST gene might be associated with cancer risk, a better understanding of chemical metabolizing gene distribution can contribute to risk assessment of humans exposed to environmental carcinogens.

Adult↗

An approach to assess ecological risk for polycyclic aromatic hydrocarbons (PAHs) in surface water from Tianjin.

Three approaches were applied and compared to evaluate additive toxic effects of eight PAHs to aquatic organisms in rivers in the Tianjin area. Although the toxicity of the studied PAH compounds did not significantly increase the risk to aquatic organisms, the results of all three approaches indicated that the additive effect of the eight PAHs was significantly stronger than any individual compound acting alone, which indicated the applicability of the approaches. Further, of the compounds studied, anthracene was the major contributor to the overall toxic effect of the mixture. The calculated geometric means of the hazard quotient for the additive effect varied from 0.00055 to 0.00062, compared to the hazard quotient of individual PAHs which ranged from 5.1 x 10-6 to 0.00053. The hazard quotient distribution geometric mean was 0.00058, with 95% of the quotient between 6.6 x 10-5 and 0.051. Overlapping areas varied from 0.00015 to 0.02 for individual PAHs and was 0.03 for additive toxicity.

Animals↗

Evaluating test methods by estimating total error.

A common procedure for evaluating a test method by comparison with another, well-accepted method has been to use a repeated measurements design, in which several individual subjects' specimens are assayed with both methods. We propose the use of the intrasubject relative mean square error, which is a function of the intrasubject relative bias and the coefficient of variation of the test method, as a measure of total error. We construct for each individual subject a score that is based on how well an individual's estimate of total error compares with a maximum allowable value. If the individual's score is > 100%, then that individual's estimate of total error exceeds the maximum allowable value. We present a distribution-free statistical methodology for evaluating the sample of scores. This involves the construction of an upper tolerance limit to determine whether the test method yields values of the total error that are acceptable for most of the population with some level of confidence. Our definition of total error is very different from that defined in the National Cholesterol Education Program (NCEP) guidelines. The NCEP bound for total error has three main problems: (a) it incorrectly assumes that the standard error of the estimated relative bias is the test coefficient of variation; (b) it incorrectly assumes that the individual estimated relative biases follow gaussian distributions; (c) it is based on requiring the relative bias of the average individual in the population to lie within prescribed limits, whereas we believe it is more important to require the total error for most of the individuals in the population, say 95%, to lie within prescribed limits.

Bias↗

Estimated global epicardial distribution of activation rate and conduction block during porcine ventricular fibrillation.

INTRODUCTION: A proposed mechanism of the maintenance of ventricular fibrillation (VF) determined by studying small hearts or segments of large hearts is that a single stable rotor exists at the site of maximal activation rate, which gives rise to activation fronts that propagate into slower activating regions where they frequently block. We wished to determine if two predictions of this hypothesized mechanism are true during VF in large hearts: (1) there is a single maximum in the distribution of activation rates with the activation rate decreasing with distance away from this maximum; and (2) the incidence of block is greater outside than inside the fastest activating region. METHODS AND RESULTS: Six 25-second episodes of VF from each of six pigs were recorded from 504 electrodes over the entire ventricular epicardium. The electrodes were divided into four zones: left ventricular base and apex (LVB and LVA) and right ventricular base and apex (RVB and RVA). A fast Fourier transform was performed on each electrogram, and the mean activation rate was estimated from the dominant (peak) frequency (DF) and block was estimated to be present during those time intervals when double peaks (DPs) were present in the power spectrum. The zones had statistically significant distributions of DF (LVB>LVA>RVA>RVB) and DP incidence (RVA>RVB>LVA>LVB). CONCLUSION: During VF, the LV base has the highest estimated activation rate and the lowest estimated block incidence, and the RV has the slowest rate but the highest block incidence. This is consistent with the concept of VF being maintained by activation fronts originating from the LV base.

Animals↗

Organization of telomeric nucleosomes: atomic force microscopy imaging and theoretical modeling.

Telomeric chromatin has peculiar features with respect to bulk chromatin, which are not fully clarified to date. Nucleosomal arrays, reconstituted on fragments of human telomeric DNA and on tandemly repeated tetramers of 5S rDNA, have been investigated at single-molecule level by atomic force microscopy and Monte Carlo simulations. A satisfactory correlation emerges between experimental and theoretical internucleosomal distance distributions. However, in the case of telomeric nucleosomal arrays containing two nucleosomes, we found significant differences. Our results show that sequence features of DNA are significant in the basic chromatin organization, but are not the only determinant.

Animals↗

Estimation of microbial cover distributions at Mammoth Hot Springs using a multiple clone library resampling method.

We propose the use of cover as a quick, low-resolution proxy for the abundance of microbial species, which reduces polymerase chain reaction bias. We showcase this concept in a computation that uses clone library information from travertine-forming hot springs in Yellowstone National Park to provide estimates of relative covers at different locations within the spring system. Samples were used from two media: the water column and the travertine substrate. The cover distribution is found to approximate a power law for samples within the water column. Significant commonality of species with the highest cover is observed in the water column for all locations, but not for species present in the substrate at different locations or between media at the same location.

Biodiversity↗

The sampling distribution of linkage disequilibrium under an infinite allele model without selection.

The sampling distributions of several statistics that measure the association of alleles on gametes (linkage disequilibrium) are estimated under a two-locus neutral infinite allele model using an efficient Monte Carlo method. An often used approximation for the mean squared linkage disequilibrium is shown to be inaccurate unless the proper statistical conditioning is used. The joint distribution of linkage disequilibrium and the allele frequencies in the sample is studied. This estimated joint distribution is sufficient for obtaining an approximate maximum likelihood estimate of C = 4Nc, where N is the population size and c is the recombination rate. It has been suggested that observations of high linkage disequilibrium might be a good basis for rejecting a neutral model in favor of a model in which natural selection maintains genetic variation. It is found that a single sample of chromosomes, examined at two loci cannot provide sufficient information for such a test if C less than 10, because with C this small, very high levels of linkage disequilibrium are not unexpected under the neutral model. In samples of size 50, it is found that, even when C is as large as 50, the distribution of linkage disequilibrium conditional on the allele frequencies is substantially different from the distribution when there is no linkage between the loci. When conditioned on the number of alleles at each locus in the sample, all of the sample statistics examined are nearly independent of theta = 4N mu, where mu is the neutral mutation rate.

Alleles↗