Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Distributions”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Seasonality in tropical AIDS: a geographical analysis.

This paper presents evidence that the growth rate of the AIDS epidemic at the district level in Uganda, Central Africa, displays a seasonally recurring geographical pattern, with epidemic acceleration in some areas of the country in the first 8 months of each year. The spatial and temporal variations in acceleration appear to be correlated with the predominant agricultural systems in different parts of Uganda. Based upon the frequently hypothesized relationship between malnourishment and the progression to clinical AIDS in HIV-infected people, it is suggested that the variations in epidemic speed reflect the seasonal patterns of nutritional deficiency which occur under some tropical agricultural systems. These preliminary findings require further verification since they have important implications for directing nutrition-related remedial responses to the AIDS epidemic in tropical countries where malnutrition and endemic HIV infection coincide.

Acquired Immunodeficiency Syndrome↗

A sella turcica bridge in subjects with dental anomalies.

Calcification of the interclinoid ligament (ICL) of the sella turcica, or sella turcica bridging, has been associated with severe craniofacial deviations. Despite no comprehensive study on the sella turcica bridge, a relationship with tooth and eruption disturbances has been reported. In order to investigate whether congenital absence of the second mandibular premolar, or the presence of a palatally displaced canine (PDC), is associated with sella bridging, a retrospective study was performed. Lateral cephalometric radiographs from 20 males and 14 females, aged between 8 and 16 years, with a PDC and second mandibular premolar aplasia were reviewed and compared with a control group. A standardized scoring scale was established to quantify the extent of a sella turcica bridge from each radiograph (no calcification, partially calcified, and completely calcified). The prevalence of complete calcification of the ICL in adolescents with dental anomalies was equal to 17.6 per cent, while an incidence 9.9 per cent was found in the control group. A partially calcified sella turcica was observed in 58.8 per cent of adolescents with dental anomalies compared with 33.7 per cent in the control group. The association between the degree of calcification of the ICL and the presence of dental anomalies in the studied adolescents was statistically significant according to chi-square statistics (P = 0.004). This was confirmed by Fisher's exact test (P = 0.003). According to these findings, the prevalence of a sella turcica bridge in adolescents with dental anomalies is increased, while age and gender do not greatly influence ossification of the ICL. The very early appearance during development of a sella turcica bridge should alert clinicians to possible tooth anomalies in life later.

Adolescent↗

On small-sample confidence intervals for parameters in discrete distributions.

The traditional definition of a confidence interval requires the coverage probability at any value of the parameter to be at least the nominal confidence level. In constructing such intervals for parameters in discrete distributions, less conservative behavior results from inverting a single two-sided test than inverting two separate one-sided tests of half the nominal level each. We illustrate for a variety of discrete problems, including interval estimation of a binomial parameter, the difference and the ratio of two binomial parameters for independent samples, and the odds ratio.

Biometry↗

Ultrasound echo envelope analysis using a homodyned K distribution signal model.

The statistics of ultrasound echo envelope signals can be used to characterize scattering media. The Rayleigh distribution and its generalized forms, the K and Rice distributions, have been previously used to model the echo signal. A more generalized statistical model, the homodyned K distribution, combines the K and Rice distribution features to better account for the statistics of the echo signal. We show that this model can give two parameters that are useful for media characterization: k, the ratio of coherent to diffuse signals, and, beta, which characterizes the clustering of scatters in the medium.

Computer Simulation↗

Estimating the emission source reduction of PM10 in central Taiwan.

Three theoretical parent frequency distributions; lognormal, Weibull and gamma were used to fit the complete set of PM10 data in central Taiwan. The gamma distribution is the best one to represent the performance of high PM10 concentrations. However, the parent distribution sometimes diverges in predicting the high PM10 concentrations. Therefore, two predicting methods, Method I: two-parameter exponential distribution and Method II: asymptotic distribution of extreme value, were used to fit the high PM10 concentration distributions more correctly. The results fitted by the two-parameter exponential distribution are better matched with the actual high PM10 data than that by the parent distributions. Both of the predicting methods can successfully predict the return period and exceedances over a critical concentration in the future year. Moreover, the estimated emission source reductions of PM10 required to meet the air quality standard by Method I and Method II are very close. The estimated emission source reductions of PM10 range from 34% to 48% in central Taiwan.

Air Pollutants↗

Body mass index, height, weight, arm circumference, and mortality in rural Bangladeshi women: a 19-y longitudinal study.

BACKGROUND: Studies in Western populations report a J- or U-shaped relation between body mass index (BMI; in kg/m(2)) and mortality, in which persons with extremes of BMI experience increased mortality. In contrast, little is known about populations in developing countries, where nutritional status is lower. OBJECTIVE: The objective was to examine the association between BMI and mortality in Bangladeshi women. DESIGN: A cohort of 1888 rural Bangladeshi women (mean age: 27.9 y) was followed over 19 y. Height, weight, arm circumference, fertility, and socioeconomic data were obtained between 1975 and 1979. Mortality, loss-to-follow-up, and additional socioeconomic data were identified by the demographic surveillance system of the International Centre for Health and Population Research, Bangladesh. Proportional hazards regression was used to examine the relation between BMI and all-cause mortality. RESULTS: The association between BMI and mortality was reverse J-shaped. After adjustment for socioeconomic indicators, the risk of dying was highest in women with BMIs in the lowest 10% of the decile distribution (< 16.39) and lowest in women with intermediate (11-89% range of the decile distribution) BMIs (16.39-20.71). Women with BMIs in the highest 10% of the distribution (> 20.71) had slightly elevated mortality (NS) compared with those with intermediate BMIs. Age and education were strongly associated with mortality. Women without schooling had a risk of mortality 4 times that of women with > or = 1 y of schooling. CONCLUSIONS: A woman's BMI relative to the BMI distribution in the local population may be a better predictor of mortality than is absolute BMI. The contribution of education in reducing mortality supports development programs aimed at increasing women's education.

Adolescent↗

Estimating haplotype-disease associations with pooled genotype data.

The genetic dissection of complex human diseases requires large-scale association studies which explore the population associations between genetic variants and disease phenotypes. DNA pooling can substantially reduce the cost of genotyping assays in these studies, and thus enables one to examine a large number of genetic variants on a large number of subjects. The availability of pooled genotype data instead of individual data poses considerable challenges in the statistical inference, especially in the haplotype-based analysis because of increased phase uncertainty. Here we present a general likelihood-based approach to making inferences about haplotype-disease associations based on possibly pooled DNA data. We consider cohort and case-control studies of unrelated subjects, and allow arbitrary and unequal pool sizes. The phenotype can be discrete or continuous, univariate or multivariate. The effects of haplotypes on disease phenotypes are formulated through flexible regression models, which allow a variety of genetic hypotheses and gene-environment interactions. We construct appropriate likelihood functions for various designs and phenotypes, accommodating Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We develop simple and efficient numerical algorithms for calculating the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. We assess the performance of the proposed methods through simulation studies, and provide an application to the Finland-United States Investigation of NIDDM Genetics Study. The results show that DNA pooling is highly efficient in studying haplotype-disease associations. As a by-product, this work provides valid and efficient methods for estimating haplotype-disease associations with unpooled DNA samples.

Algorithms↗

Gene-centromere distances of allozyme loci in even- and odd-year pink salmon, (Oncorhynchus gorbuscha).

We produced gynogenetic progeny families to estimate gene-centromere (G-C) distances of allozyme loci in even-year and odd-year pink salmon (Oncorhynchus gorbuscha). G-C distances of 37 loci distributed on a chromosome ranged from 1 cM at LDH-A1* to 49 cM at ADA-2*, DIA-2*, and sMDH-B1,2*. The distribution of the G-C distances along the chromosome arm was not even and appears telomeric. Eight loci in even-year and seven in odd-year showed high G-C distances (>45 cM), indicating that one crossover per chromosome arm is usual in pink salmon. Variation was observed in the results from different families; 14 loci out of 21 tested, showed heterogeneity. At mAH-3*, G-C distances from five odd-year families ranged from 6 to 37 cM; the widest range observed in this study. At isoloci such as sMDH-A 1,2* and sMDH-B1,2* the distances from different families were grouped into statistically discrete distributions, suggesting that it may be a reflection polymorphism at both isoloci. It appears G-C distances in salmonid species are well conserved with some minor differences.

Animals↗

A Markov chain model for animal estrous cycling data.

Estrous cycling data contain sequences of characters (e.g., DPEMD). Each sequence represents an animal's estrous cycle, with each character indicating the daily estrous cycle stage. Changes in the estrous cycle pattern, which is determined by estrous stage lengths, can provide information on adverse events. Stage lengths are not directly observable. However interval censored lengths for all but the first and the last stages in a sequence can be extracted from the data. We propose a Markov chain model to approximate the estrous cycling process. The transition probabilities from one stage to another can be derived by conditioning on stage lengths. Assuming Weibull distribution for stage lengths, with the second Weibull parameter depending upon treatment effects and animal-specific random effects, regression models on censored stage lengths are fitted. A Bayesian approach is used for inference on dose effects. The analysis is implemented with MCMC method in WinBUGS. An estrous cycling data set from a National Toxicology Program study is analyzed as an example.

Animals↗

Fertility and adaptation: Indochinese refugees in the United States.

"Levels of fertility among Indochinese refugees in the United States are explored in the context of a highly compressed demographic transition implicit in the move from high-fertility Southeast Asian societies to a low-fertility resettlement region. A theoretical model is developed to explain the effect on refugee fertility of social background characteristics, migration history and patterns of adaptation to a different economic and cultural environment controlling for marital history and length of residence in the U.S." The chief source for the data and analyses is the Indochinese Health and Adaptation Research Project (IHARP), San Diego State University. "Multiple regression techniques are used to test the model which was found to account for nearly half of the variation in refugee fertility levels in the United States. Fertility is much higher for all Indochinese ethnic groups than it is for American women; the number of children in refugee families is in turn a major determinant of welfare dependency. Adjustments for rates of natural increase indicate a total 1985 Indochinese population of over one million, making it one of the largest Asian-origin populations in the United States."

Acculturation↗

Distribution of protein folds in the three superkingdoms of life.

A sensitive protein-fold recognition procedure was developed on the basis of iterative database search using the PSI-BLAST program. A collection of 1193 position-dependent weight matrices that can be used as fold identifiers was produced. In the completely sequenced genomes, folds could be automatically identified for 20%-30% of the proteins, with 3%-6% more detectable by additional analysis of conserved motifs. The distribution of the most common folds is very similar in bacteria and archaea but distinct in eukaryotes. Within the bacteria, this distribution differs between parasitic and free-living species. In all analyzed genomes, the P-loop NTPases are the most abundant fold. In bacteria and archaea, the next most common folds are ferredoxin-like domains, TIM-barrels, and methyltransferases, whereas in eukaryotes, the second to fourth places belong to protein kinases, beta-propellers and TIM-barrels. The observed diversity of protein folds in different proteomes is approximately twice as high as it would be expected from a simple stochastic model describing a proteome as a finite sample from an infinite pool of proteins with an exponential distribution of the fold fractions. Distribution of the number of domains with different folds in one protein fits the geometric model, which is compatible with the evolution of multidomain proteins by random combination of domains. [Fold predictions for proteins from 14 proteomes are available on the World Wide Web at. The FIDs are available by anonymous ftp at the same location.]

Algorithms↗

Lorenz curves and their use in describing the distribution of 'the total burden' of dental caries in a population.

PURPOSE: 1) to describe the distribution of the total burden of dental caries in Danish adolescents over a 15-year period using Lorenz curves and, 2) to compare the observed distributions with Poisson distributions. METHOD: caries data for 15-year-old adolescents reported to the database for the national reporting system for the Danish Municipal Dental Service for Children and Adolescents in 1980 (n = 61,621) and 1995 (n = 50,359). RESULTS: The DMFS cut-off point for a given percentile had decreased from 1980 to 1995 and Lorenz curves showed a pattern of increasing inequality, even when only diseased individuals (i.e. with DMFS > or = 1) were included. The dispersion was larger than could be expected, if caries developed according to a random pattern modelled by e.g. a Poisson distribution. CONCLUSIONS: Lorenz curves may be a useful tool in the analysis of caries data, with special reference to determining the appropriateness of implementing high-risk preventive strategies.

Adolescent↗

Estimation of erythrocyte population state by the spherical index distribution.

The densities of cell distributions by spherical index (SI) in erythrocyte populations from healthy adults and donors with endocrine pathologies were determined via the developed method. The investigation shows that this characteristic varies for different donors, thereby reflecting the erythrocyte population state of an individual donor. Individual distribution curves obtained from healthy donors are close to Gaussian and are characterized by smooth curve plot with one maximum. Cells distribution by SI in donors with endocrine pathologies has a polymodal character. Our research shows that the developed method for determining erythrocyte distribution density by SI is a sensitive and informative test for quantitative evaluation of an erythrocyte population state. Moreover, this characteristic has clear physical and physiological significance, because an erythrocyte shape is strongly conditioned by the cell age and influences the ability to pass through microcapillaries in blood circulation.

Case-Control Studies↗

The homocysteine distribution: (mis)judging the burden.

The nonfasting plasma total homocysteine (P-tHcy) concentration was measured in a random sample of 3025 Dutch adults aged 20-65 years (main study). The positively skewed distribution had a geometric mean of 13.9 micromol/L in men and 12.6 micromol/L in women. Blood of the main study was not cooled or centrifuged immediately after drawing. A stability study (n = 26) indicated that this could have resulted in a small (0.4 micromol/L) overestimation of the means. A comparative study (n = 88), and a reproduction of these results in an entirely different population (n = 213), showed a systematic difference in P-tHcy concentration of -2.4 micromol/L between our laboratory (Nijmegen, the Netherlands) and that in Bergen, Norway. With the information of the additional studies we provided precise and valid data of the Dutch P-tHcy distribution, from which we conclude the status in the Netherlands is worse than in other European countries. Furthermore, we showed that comparison of P-tHcy data is complicated unless the interlaboratory differences are known. @ 2001 Elsevier Science Inc.

Adult↗

The insertional history of an active family of L1 retrotransposons in humans.

As humans contain a currently active L1 (LINE-1) non-LTR retrotransposon family (Ta-1), the human genome database likely provides only a partial picture of Ta-1-generated diversity. We used a non-biased method to clone Ta-1 retrotransposon-containing loci from representatives of four ethnic populations. We obtained 277 distinct Ta-1 loci and identified an additional 67 loci in the human genome database. This collection represents approximately 90% of the Ta-1 population in the individuals examined and is thus more representative of the insertional history of Ta-1 than the human genome database, which lacked approximately 40% of our cloned Ta-1 elements. As both polymorphic and fixed Ta-1 elements are as abundant in the GC-poor genomic regions as in ancestral L1 elements, the enrichment of L1 elements in GC-poor areas is likely due to insertional bias rather than selection. Although the chromosomal distribution of Ta-1 inserts is generally a function of chromosomal length and gene density, chromosome 4 significantly deviates from this pattern and has been much more hospitable to Ta-1 insertions than any other chromosome. Also, the intra-chromosomal distribution of Ta-1 elements is not uniform. Ta-1 elements tend to cluster, and the maximal gaps between Ta-1 inserts are larger than would be expected from a model of uniform random insertion.

Chromosome Mapping↗

General schema theory for genetic programming with subtree-swapping crossover: part I.

This is the first part of a two-part paper which introduces a general schema theory for genetic programming (GP) with subtree-swapping crossover. The theory is based on a Cartesian node reference system which makes it possible to describe programs as functions over the space N(2) and allows one to model the process of selection of the crossover points of subtree-swapping crossovers as a probability distribution over N(4). In Part I, we present these notions and models and show how they can be used to calculate useful quantities. In Part II we will show how this machinery, when integrated with other definitions, such as that of variable-arity hyperschema, can be used to construct a general and exact schema theory for the most commonly used types of GP.

Algorithms↗