Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Some statistical aspects of food intake assessment.

OBJECTIVE: To present the results of the statistical working group of the EFCOSUM project on estimating the minimum sample size for a pan-European dietary survey. BACKGROUND AND METHODS: Numerous statistical issues are involved when planning a nutritional survey aimed at evaluating various indicators, especially if it will be carried out in different countries. The plenary workshop of the EFCOSUM project has chosen four relevant statistical topics: the sample size estimation for dietary surveys, the number of repeated measurements needed to estimate usual intake for each individual; the statistical presentation of data; and the statistical procedures for estimating the usual intake distribution from a limited number of days of observation. This article deals with the first three topics mentioned. The participants of the EFCOSUM project answered a small questionnaire in order to get agreement on the method of estimating a minimum sample size in the context of a monitoring of dietary indicators. Data on the variability of dietary indicators of interest was also collected, in order to calculate a minimum sample size. RESULTS AND CONCLUSION: The main result was that a minimum sample size of 2000 adults in each European country will be needed in order to identify trends in the mean intake of the most relevant foods and nutrients in Europe. This sample size should be higher if trends have to be indentified for socio-demographic subgroups.

Data Interpretation, Statistical↗

Redressing the power and effect of significance. A new approach to an old problem: teaching statistics to nursing students.

Many barriers to learning are present when teaching research methods. Developing, within students of nursing, the skills of reading and interpreting research reports is vital if the profession is to contribute to the general aim of achieving a sound basis for all health care interventions. This paper overviews the current move toward evidence based practice, the challenges that are present when teaching research to nursing students and offers an approach to teaching quantitative research that will help students of nursing to understand the key concepts that form the basis of inferential statistics. In this work we argue that the traditional emphasis on probability and statistical significance needs to be redressed and that effect size and power should form the basis of teaching students the concepts involved in inferential statistics. We argue that introducing students to the key concepts in statistical decision making in a particular order, effect size then power and lastly statistical significance, will lead to a better understanding of Type I and Type II errors. After all, the purpose of hypothesis testing is to detect a treatment or intervention effect. Power is dependent upon the size of the treatment effect, thus it must be introduced after effect size. Students, we argue, must be able to understand the concept of effect size. We consider this to be a foundational concept that will help to develop a firmer grasp of the decision making processes involved in hypothesis testing. Such an approach will form a more logical approach to teaching this subject and will allow for the use of real world examples to form the basis of learning.

Bias↗

Design and statistical methods in studies using animal models of development.

Experiments involving neonates should follow the same basic principles as most other experiments. They should be unbiased, be powerful, have a good range of applicability, not be excessively complex, and be statistically analyzable to show the range of uncertainty in the conclusions. However, investigation of growth and development in neonatal multiparous animals poses special problems associated with the choice of "experimental unit" and differences between litters: the "litter effect." Two main types of experiments are described, with recommendations regarding their design and statistical analysis: First, the "between litter design" is used when females or whole litters are assigned to a treatment group. In this case the litter, rather than the individuals within a litter, is the experimental unit and should be the unit for the statistical analysis. Measurements made on individual neonatal animals need to be combined within each litter. Counting each neonate as a separate observation may lead to incorrect conclusions. The number of observations for each outcome ("n") is based on the number of treated females or whole litters. Where litter sizes vary, it may be necessary to use a weighted statistical analysis because means based on more observations are more reliable than those based on a few observations. Second, the more powerful "within-litter design" is used when neonates can be individually assigned to treatment groups so that individuals within a litter can have different treatments. In this case, the individual neonate is the experimental unit, and "n" is based on the number of individual pups, not on the number of whole litters. However, variation in litter size means that it may be difficult to perform balanced experiments with equal numbers of animals in each treatment group within each litter. This increases the complexity of the statistical analysis. A numerical example using a general linear model analysis of variance is provided in the Appendix. The use of isogenic strains should be considered in neonatal research. These strains are like immortal clones of genetically identical individuals (i.e., they are uniform, stable, and repeatable), and their use should result in more powerful experiments. Inbred females mated to males of a different inbred strain will produce F1 hybrid offspring that will be uniform, vigorous, and genetically identical. Different strains may develop at different rates and respond differently to experimental treatments.

Animals↗

Protein sequence-structure compatibility criteria in terms of statistical hypothesis testing.

The assignment of query protein sequences to probable folds in a threading approach is based on the statistical analysis (learning) of structural properties of amino acids in known protein structures. We formalize the recognition problem in terms of mathematical statistics, namely statistical hypothesis testing. Our general formulation leads to various mathematical forms of a decision rule function for evaluation of the quality of a sequence-structure fit. Three criteria were derived according to a likelihood ratio approach. Two of them have new functional forms while the third happens to coincide with the mean force potential function previously derived under the additional assumption of the Boltzmann law. New decision rule functions employ (i) the Parzen estimator of a probability density and (ii) the newly introduced non-parametric statistic with known asymptotic distribution. We compared criteria efficiency by a 'structure seeks sequence' search for three highly populated template folds through a query library of non-homologous sequences of proteins with known 3D structure using residue accessibility as an environmental variable. Various criteria reflect different underlying statistical propositions and thus often recognize diverse correct sequence-structure matches. On the other hand, if an amino acid sequence is recognized as compatible with a template by each of three decision rules it appears that one can make a more reliable inference of sequence-structure relationship since almost all false positives obtained by the three criteria differ.

Algorithms↗

Statistical maps for EEG dipolar source localization.

We present a method that estimates three-dimensional statistical maps for electroencephalogram (EEG) source localization. The maps assess the likelihood that a point in the brain contains a dipolar source, under the hypothesis of one, two or three activated sources. This is achieved by examining all combinations of one to three dipoles on a coarse grid and attributing to each combination a score based on an F statistic. The probability density function of the statistic under the null hypothesis is estimated nonparametrically, using bootstrap resampling. A theoretical F distribution is then fitted to the empirical distribution in order to allow correction for multiple comparisons. The maps allow for the systematic exploration of the solution space for dipolar sources. They permit to test whether the data support a given solution. They do not rely on the assumption of uncorrelated source time courses. They can be compared to other statistical parametric maps such as those used in functional magnetic resonance imaging (fMRI). Results are presented for both simulated and real data. The maps were compared with LORETA and MUSIC results. For the real data consisting of an average of epileptic spikes, we observed good agreement between the EEG statistical maps, intracranial EEG recordings, and fMRI activations.

Brain↗

An improved statistical methodology to estimate and analyze impedances and transfer functions.

Estimating the mathematical relationship between pulsatile time series (e.g., pressure and flow) is an effective technique for studying dynamic systems. The frequency-domain relationship between time series, often calculated as an impedance (pressure/flow), is known more generally as a frequency-response or transfer function (output/input). Current statistical methods for transfer function analysis 1) assume erroneously that repeated observations on a subject are independent, 2) have limited statistical value and power, or 3) are restricted to use in single subjects rather than in an entire sample. This paper develops a regression model for transfer function analysis that corrects each of these deficiencies. Spectral densities of the input and output time series and the cross-spectral density between them are first estimated from discrete Fourier transforms and then used to obtain regression estimates of the transfer function. Statistical comparisons of the transfer function estimates use a test statistic that is distributed as chi2. Confidence intervals for amplitude and phase can also be calculated. By correctly modeling repeated observations on each subject, this improved statistical approach to transfer function estimation and analysis permits the simultaneous analysis of data from all subjects in a sample, improves the power of the transfer function model, and has broad relevance to the study of dynamic physiological systems.

Blood Pressure↗

Utilization of two sample t-test statistics from redundant probe sets to evaluate different probe set algorithms in GeneChip studies.

BACKGROUND: The choice of probe set algorithms for expression summary in a GeneChip study has a great impact on subsequent gene expression data analysis. Spiked-in cRNAs with known concentration are often used to assess the relative performance of probe set algorithms. Given the fact that the spiked-in cRNAs do not represent endogenously expressed genes in experiments, it becomes increasingly important to have methods to study whether a particular probe set algorithm is more appropriate for a specific dataset, without using such external reference data. RESULTS: We propose the use of the probe set redundancy feature for evaluating the performance of probe set algorithms, and have presented three approaches for analyzing data variance and result bias using two sample t-test statistics from redundant probe sets. These approaches are as follows: 1) analyzing redundant probe set variance based on t-statistic rank order, 2) computing correlation of t-statistics between redundant probe sets, and 3) analyzing the co-occurrence of replicate redundant probe sets representing differentially expressed genes. We applied these approaches to expression summary data generated from three datasets utilizing individual probe set algorithms of MAS5.0, dChip, or RMA. We also utilized combinations of options from the three probe set algorithms. We found that results from the three approaches were similar within each individual expression summary dataset, and were also in good agreement with previously reported findings by others. We also demonstrate the validity of our findings by independent experimental methods. CONCLUSION: All three proposed approaches allowed us to assess the performance of probe set algorithms using the probe set redundancy feature. The analyses of redundant probe set variance based on t-statistic rank order and correlation of t-statistics between redundant probe sets provide useful tools for data variance analysis, and the co-occurrence of replicate redundant probe sets representing differentially expressed genes allows estimation of result bias. The results also suggest that individual probe set algorithms have dataset-specific performance.

Algorithms↗

Spatial statistical methods in health.

The study of the geographical distribution of disease incidence and its relationship to potential risk factors (referred to here as "geographical epidemiology") has provided, and continues to provide, rich ground for the application and development of statistical methods and models. In recent years increasingly powerful and versatile statistical tools have been developed in this application area. This paper discusses the general classes of problem in geographical epidemiology and reviews the key statistical methods now being employed in each of the application areas identified. The paper does not attempt to exhaustively cover all possible methods and models, but extensive references are provided to further details and to additional approaches. The overall aim is to provide a picture of the "current state of the art" in the use of spatial statistical methods in epidemiological and public health research. Following the review of methods, the main software environments which are available to implement such methods are discussed. The paper concludes with some brief general reflections on the epidemiological and public health implications of the use of spatial statistical methods in health and on associated benefits and problems.

Cluster Analysis↗

Statistics for correlated data: phylogenies, space, and time.

Here we give an introduction to the growing number of statistical techniques for analyzing data that are not independent realizations of the same sampling process--in other words, correlated data. We focus on regression problems, in which the value of a given variable depends linearly on the value of another variable. To illustrate different types of processes leading to correlated data, we analyze four simulated examples representing diverse problems arising in ecological studies. The first example is a comparison among species to determine the relationship between home-range area and body size; because species are phylogenetically related, they do not represent independent samples. The second example addresses spatial variation in net primary production and how this might be affected by soil nitrogen; because nearby locations are likely to have similar net primary productivity for reasons other than soil nitrogen, spatial correlation is likely. In the third example, we consider a time-series model to ask whether the decrease in density of a butterfly species is the result of decreases in its host-plant density; because the population density of a species in one generation is likely to affect the density in the following generation, time-series data are often correlated. The fourth example combines both spatial and temporal correlation in an experiment in which prey densities are manipulated to determine the response of predators to their food supply. For each of these examples, we use a different statistical approach for analyzing models of correlated data. Our goal is to give an overview of conceptual issues surrounding correlated data, rather than a detailed tutorial in how to apply different statistical techniques. By dispelling some of the mystery behind correlated data, we hope to encourage ecologists to learn about statistics that could be useful in their own work. Although at first encounter these techniques might seem complicated, they have the power to simplify ecological research by making more types of data and experimental designs open to statistical evaluation.

Animals↗

Using the open-source statistical language R to analyze the dichotomous Rasch model.

R, an open-source statistical language and data analysis tool, is gaining popularity among psychologists currently teaching statistics. R is especially suitable for teaching advanced topics, such as fitting the dichotomous Rasch model--a topic that involves transforming complicated mathematical formulas into statistical computations. This article describes R's use as a teaching tool and a data analysis software program in the analysis of the Rasch model in item response theory. It also explains thetheory behind, as well as an educator's goals for, fitting the Rasch model with joint maximum likelihood estimation. This article also summarizes the R syntax for parameter estimation and the calculation of fit statistics. The results produced by R is compared with the results obtained from MINISTEP and the output of a conditional logit model. The use of R is encouraged because it is free, supported by a network of peer researchers, and covers both basic and advanced topics in statistics frequently used by psychologists.

Algorithms↗

Statistical analysis of the seasonal variation in demographic data.

There has been little agreement as to whether reproduction or similar demographic events occur seasonally and, especially, whether there is any universal seasonal pattern. One reason is that the seasonal pattern may vary in different populations and at different times. Another reason is that different statistical methods have been used. Every statistical model is based on certain assumed conditions and hence is designed to identify specific components of the seasonal pattern. Therefore, the statistical method applied should be chosen with due consideration. In this study we present, develop, and compare different statistical methods for the study of seasonal variation. Furthermore, we stress that the methods are applicable for the analysis of many kinds of demographic data. The first approaches in the literature were based on monthly frequencies, on the simple sine curve, and on the approximation that the months are of equal length. Later, "the population at risk" and the fact that the months have different lengths were considered. Under these later assumptions the targets of the statistical analyses are the rates. In this study we present and generalize the earlier models. Furthermore, we use trigonometric regression methods. The trigonometric regression model in its simplest form corresponds to the sine curve. We compare the regression methods with the earlier models and reanalyze some data. Our results show that models for rates eliminate the disturbing effects of the varying length of the months, including the effect of leap years, and of the seasonal pattern of the population at risk. Therefore, they give the purest analysis of the seasonal pattern of the demographic data in question, e.g., rates of general births, twin maternities, neural tube defects, and mortality. Our main finding is that the trigonometric regression methods are more flexible and easier to handle than the earlier methods, particularly when the data differ from the simple sine curve.

Algorithms↗

Domain analysis and modeling to improve comparability of health statistics.

Health statistics is an essential element to improve the ability of managers of health institutions, healthcare researchers, policy makers, and health professionals to formulate appropriate course of reactions and to make decisions based on evidence. To ensure adequate health statistics, standards are of critical importance. A study on healthcare statistics domain analysis is underway in an effort to improve usability and comparability of health statistics. The ongoing study focuses on structuring the domain knowledge and making the knowledge explicit with a data element dictionary being the core. Supplemental to the dictionary are a domain term list, a terminology dictionary, and a data model to help organize the concepts constituting the health statistics domain.

Demography↗

Vital statistics in the United States: preparing for the next century.

"This paper outlines the development of U.S. national vital statistics based on the local registration of vital events in the United States during the twentieth century, including the organization of the National Vital Statistics System. Current data developments and selected publications of the National Center for Health Statistics are presented as they relate to vital statistics. The paper concludes with an overview of ongoing efforts at the local, state, and federal levels to improve the timeliness and quality of vital statistics through the redesign and automation of data collection, processing, and dissemination systems."

Americas↗

Statistical hypothesis testing--how exact are exact p-values?

OBJECTIVES AND BACKGROUND: When testing a hypothesis statistically, a principle is generally accepted that exact p values shall be stated in the treatise. Researchers have the choice of many statistical computer programmes with implemented hypothesis tests. Are exact p values calculated in the same statistical tests by diverse statistical programmes identical? METHODS: The respective zero hypothesis were tested in 5 artificially created data sets by the parametric unpaired t-test, non-parametric Mann-Whitney test, two-tailed F-test. The calculations were carried out by the following programmes: Statistix, version 7.1 (source www.statistix.com), Analyse-it, version 1.62 (source www.analyse-it.com), MedCalc, version 6.14 (source www.medcalc.be). The p values in the same tests were mutually compared. RESULTS: All three programmes calculated identical exact p values for the t-test. In the remaining two tests in case of 26 out of 44 calculations (59.1 per cent; 95 per cent confidence interval 43-73 per cent) different p values were calculated. The greatest difference was 18.35 per cent. In two cases the values oscillated about 0.05 and this fact caused essentially different interpretation of results. CONCLUSIONS: Using the significance test in the biomedical research has been subject to criticism for a longer period of time. The testing of the zero hypothesis on the arbitrary significance level of 0.05 should be substituted by other methods. Our discoveries should undermine the ungrounded belief of the users of statistical tests--physicians in ununderminable accuracy of mathematical procedures. The use of confidence intervals deems much more suitable although there are objections against them as well. (Tab. 4, Fig. 1, Ref. 19.).

Confidence Intervals↗

Securing cooperation from persons supplying statistical data.

Securing the co-operation of persons supplying information required for medical statistics is essentially a problem in human relations, and an understanding of the motivations, attitudes, and behaviour of the respondents is necessary.Before any new statistical survey is undertaken, it is suggested by Aubenque and Harris that a preliminary review be made so that the maximum use is made of existing information. Care should also be taken not to burden respondents with an overloaded questionnaire. Aubenque and Harris recommend simplified reporting. Complete population coverage is not necessary.Neurdenburg suggests that the co-operation and support of such organizations as medical associations and social security boards are important and that propaganda should be directed specifically to the groups whose co-operation is sought. Informal personal contacts are valuable and desirable, according to Blaikley, but may have adverse effects if the right kind of approach is not made.Financial payments as an incentive in securing co-operation are opposed by Neurdenburg, who proposes that only postage-free envelopes or similar small favours be granted. Blaikley and Harris, on the other hand, express the view that financial incentives may do much to gain the support of those required to furnish data; there are, however, other incentives, and full use should be made of the natural inclinations of respondents. Compulsion may be necessary in certain instances, but administrative rather than statutory measures should be adopted. Penalties, according to Aubenque, should be inflicted only when justified by imperative health requirements.The results of surveys should be made available as soon as possible to those who co-operated, and Aubenque and Harris point out that they should also be of practical value to the suppliers of the information.Greater co-operation can be secured from medical persons who have an understanding of the statistical principles involved; Aubenque and Neurdenburg suggest that a course in elementary statistical methodology be introduced in the curriculum of medical schools.Methods for improving co-operation, and thus the compilation of statistics, in particular countries are discussed by Lal and de Shelly Hernández, who deal with India and Venezuela respectively.

Data Collection↗

Survey of statistical methods used in the veterinary medical literature.

Articles published in 1992 in 6 veterinary journals were reviewed. In 51% of the articles, statistical analyses were not performed or only descriptive statistics (eg, mean, median, standard deviation) were used. The most commonly used statistical tests were ANOVA and t-tests. Knowledge of 5 categories of statistical methods (ANOVA, t-tests, contingency tables, nonparametric tests, and simple linear regression) permitted access to 90% of the veterinary literature surveyed. These data may be useful when modifying the veterinary curriculum to reflect current statistical usage.

Analysis of Variance↗

On lies and health statistics: some Latin American examples.

New methods of demographic analysis are producing estimates of fertility and mortality which are sometimes at great variance with "official" figures generated by the statistics organizations of the different countries and which are reproduced in international reference books. This discrepancy is greatest with regard to infant mortality. Using Latin American examples, the magnitude of this discrepancy is explored, biases in estimating causes of mortality are identified, and a consideration is made of morbidity figures, which, as they are generated by health care systems with very low coverages of population, tend to seriously underrepresent the prevalent levels of disease. A structural interpretaton is made of the Latin American situation, linking this crisis of health statistics with a more general crisis of the "developmentist" model under which these systems flourished, and with an upsurge in political repression in the Continent which will tend in future to increase the inaccuracy of "official" health statistics data. Finally, alternative health statistics procedures are proposed.

Child, Preschool↗

Population and vital statistics, 1981.

"For various reasons some of the data relating to population estimates, vital statistics and causes of death in 1981 were not included in the Statistical Abstract of Israel No. 33, 1982. The purpose of this [article] is to complete the missing data and to revise and update some other data." Statistics are included on population by age, sex, marital status, population group, origin, continent of birth, period of immigration, and religion; marriages, divorces, live births, deaths, natural increase, infant deaths, and stillbirths by religion; characteristics of persons marrying and divorcing, including place of residence, religion, age, previous marital status, and year and duration of marriage; live births, deaths, and infant deaths by district, sub-district, and type of locality of residence; deaths by age, sex, and continent of birth; infant deaths by age, sex, and population group; and selected life table values by population group and sex.

Age Distribution↗