Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Statistical representation and simulation of high-dimensional deformations: application to synthesizing brain deformations.

This paper proposes an approach to effectively representing the statistics of high-dimensional deformations, when relatively few training samples are available, and conventional methods, like PCA, fail due to insufficient training. Based on previous work on scale-space decomposition of deformation fields, herein we represent the space of "valid deformations" as the intersection of three subspaces: one that satisfies constraints on deformations themselves, one that satisfies constraints on Jacobian determinants of deformations, and one that represents smooth deformations via a Markov Random Field (MRF). The first two are extensions of PCA-based statistical shape models. They are based on a wavelet packet basis decomposition that allows for more accurate estimation of the covariance structure of deformation or Jacobian fields, and they are used jointly due to their complementary strengths and limitations. The third is a nested MRF regularization aiming at eliminating potential discontinuities introduced by assumptions in the statistical models. A randomly sampled deformation field is projected onto the space of valid deformations via iterative projections on each of these subspaces until convergence, i.e. all three constraints are met. A deformation field simulator uses this process to generate random samples of deformation fields that are not only realistic but also representative of the full range of anatomical variability. These simulated deformations can be used for validation of deformable registration methods. Other potential uses of this approach include representation of shape priors in statistical shape models as well as various estimation and hypothesis testing paradigms in the general fields of computational anatomy and pattern recognition.

Algorithms↗

The kappa statistic in rehabilitation research: an examination.

The number and sophistication of statistical procedures reported in medical rehabilitation research is increasing. Application of the principles and methods associated with evidence-based practice has contributed to the need for rehabilitation practitioners to understand quantitative methods in published articles. Outcomes measurement and determination of reliability are areas that have experienced rapid change during the past decade. In this study, distinctions between reliability and agreement are examined. Information is presented on analytical approaches for addressing reliability and agreement with the focus on the application of the kappa statistic. The following assumptions are discussed: (1) kappa should be used with data measured on a categorical scale, (2) the patients or objects categorized should be independent, and (3) the observers or raters must make their measurement decisions and judgments independently. Several issues related to using kappa in measurement studies are described, including use of weighted kappa, methods of reporting kappa, the effect of bias and prevalence on kappa, and sample size and power requirements for kappa. The kappa statistic is useful for assessing agreement among raters, and it is being used more frequently in rehabilitation research. Correct interpretation of the kappa statistic depends on meeting the required assumptions and accurate reporting.

Activities of Daily Living↗

Statistical power and estimation of the number of required subjects for a study based on the t-test: a surgeon's primer.

The underlying concepts for calculating the power of a statistical test elude most investigators. Understanding them helps to know how the various factors contributing to statistical power factor into study design when calculating the required number of subjects to enter into a study. Most journals and funding agencies now require a justification for the number of subjects enrolled into a study and investigators must present the principals of powers calculations used to justify these numbers. For these reasons, knowing how statistical power is determined is essential for researchers in the modern era. The number of subjects required for study entry, depends on the following four concepts: 1) The magnitude of the hypothesized effect (i.e., how far apart the two sample means are expected to differ by); 2) the underlying variability of the outcomes measured (standard deviation); 3) the level of significance desired (e.g., alpha = 0.05); 4) the amount of power desired (typically 0.8). If the sample standard deviations are small or the means are expected to be very different then smaller numbers of subjects are required to ensure avoidance of type 1 and 2 errors. This review provides the derivation of the sample size equation for continuous variables when the statistical analysis will be the Student's t-test. We also provide graphical illustrations of how and why these equations are derived.

Data Interpretation, Statistical↗

Suggested statistical standards for NTT manuscripts: notes from two of your reviewers.

The purpose of a scientific paper in this journal is to persuade the reader of some important or potentially important facts. For a reader to be persuaded, first the manuscript reviewers must be persuaded, and if the manuscript involves statistical reasoning, at least one of those reviewers is likely to be a statistician. This invited article, by two long-time reviewers for Neurotoxicology and Teratology (NTT) who are also contributors of statistical papers, surveys some of the principles that render a manuscript more persuasive or less persuasive in our eyes. These principles are overwhelmingly not statistical but logical. For one typical NTT manuscript theme, the relation between some toxic exposure and one or more negative outcomes in humans, the aspects of manuscripts we scrutinize most closely include biological plausibility, dose-response relationships, breadth of evidence, adjustments for measurement bias, attention to assumptions and scatterplots in the search for confounds, and, in general, a sincere attempt to enunciate and then refute plausible hypotheses rival to the one the investigators prefer. The literature of excellent studies in other fields provides ample instances of good practice in these matters; we review it in those fields for applications in ours. Formal statistical significance testing plays almost no role in the most persuasive papers. In particular, findings that appear only after "adjustment for covariates" are never considered credible by these reviewers; we explain our reasons at length, and suggest alternatives.

Bias↗

Statistical analysis and biological interpretation of the flow cytometric heterogeneity observed in bacterial axenic cultures.

Histogram comparison and meaningful statistics in flow cytometry is probably the most widely encountered mathematical problem in flow cytometry. Ideally, a test for determining the statistical equality or difference of flow cytometric distributions will identify the significant differences or similarities of the obtained histograms. This situation is of particular interest when flow cytometry is used to study the heterogeneity of axenic bacterial populations. We have statistically measured the heterogeneity of successive cytometric measures, the modifications produced after 20 transfers from the same culture, and the differences between 20 subcultures of identical origin. The heterogeneity of the bacterial populations and the similarity of the obtained 360 histograms were analysed by standard statistical methods. We have studied bacterial axenic cultures in order to detect, quantify and interpret their cytometric heterogeneity, and to assess intrinsic differences and differences produced by laboratory manipulations. We concluded that the standard axenic cultures have a considerable intrinsic cellular and molecular heterogeneity. We suggest that the heterogeneity we have detected basically has two origins: cell size diversity and cell cycle variations.

Bacteria↗

Statistical approaches to experimental design and data analysis of in vivo studies.

The objective of any experiment is to obtain an unbiased and precise estimate of a treatment effect in an efficient manner. Statistical aspects of the design, conduct, and analysis of the experiment play a major role in determining whether this goal is met. We highlight some of the more important statistical issues that pertain to in vivo studies. Particular emphasis is placed on the role of randomization, the number of animals, the utilization of repeated measures data, adjustments for missing data, and dealing with multiple causes of death or treatment failure. The discussion is not intended to be a comprehensive guide to all the statistical issues that can occur in animal experiments. Rather, the objective is to acquaint researchers with components of the experiment that will require careful statistical thought.

Animals↗

When are people persuaded by DNA match statistics?

The way in which statistical DNA evidence is presented to legal decision makers can have a profound impact on the persuasiveness of that evidence. Evidence that is presented one way may convince most people that the suspect is almost certainly the source of DNA evidence recovered from a crime scene. However, when the evidence is presented another way, a sizable minority of people equally convinced that the suspect is almost certainly not the source of the evidence. Three experiments are presented within the context of a theory ("exemplar cueing theory") for when people will find statistical match evidence to be more and less persuasive. The theory holds that the perceived probative value of statistical match evidence depends on the cognitive availability of coincidental match exemplars. When legal decision makers find it hard to imagine others who might match by chance, the evidence will seem compelling. When match exemplars are readily available, the evidence will seem less compelling. Experiments 1 and 2 show that DNA match statistics that target the individual suspect and that are framed as probabilities (i.e., "The probability that the suspect would match the blood drops if he were not their source is 0.1%") are more persuasive than mathematically equivalent presentations that target a broader reference group and that are framed as frequencies ("One in 1,000 people in Houston would also match the blood drops"). Experiment 3 shows that the observed effects are less likely to occur at extremely small incidence rates. Implications for the strategic use of presentation effects at trial are considered.

Adult↗

Advances in statistical methods for substance abuse prevention research.

The paper describes advances in statistical methods for prevention research with a particular focus on substance abuse prevention. Standard analysis methods are extended to the typical research designs and characteristics of the data collected in prevention research. Prevention research often includes longitudinal measurement, clustering of data in units such as schools or clinics, missing data, and categorical as well as continuous outcome variables. Statistical methods to handle these features of prevention data are outlined. Developments in mediation, moderation, and implementation analysis allow for the extraction of more detailed information from a prevention study. Advancements in the interpretation of prevention research results include more widespread calculation of effect size and statistical power, the use of confidence intervals as well as hypothesis testing, detailed causal analysis of research findings, and meta-analysis. The increased availability of statistical software has contributed greatly to the use of new methods in prevention research. It is likely that the Internet will continue to stimulate the development and application of new methods.

Humans↗

Statistical methods for longitudinal research on bipolar disorders.

OBJECTIVES: Outcomes research in bipolar disorders, because of complex clinical variation over-time, offers demanding research design and statistical challenges. Longitudinal studies involving relatively large samples, with outcome measures obtained repeatedly over-time, are required. In this report, statistical methods appropriate for such research are reviewed. METHODS: Analytic methods appropriate for repeated measures data include: (i) endpoint analysis; (ii) endpoint analysis with last observation carried forward; (iii) summary statistic methods yielding one summary measure per subject; (iv) random effects and generalized estimating equation (GEE) regression modeling methods; and (v) time-to-event survival analyses. RESULTS: Use and limitations of these several methods are illustrated within a randomly selected (33%) subset of data obtained in two recently completed randomized, double blind studies on acute mania. Outcome measures obtained repeatedly over 3 or 4 weeks of blinded treatment in active drug and placebo sub-groups included change-from-baseline Young Mania Rating Scale (YMRS) scores (continuous measure) and achievement of a clinical response criterion (50% YMRS reduction). Four of the methods reviewed are especially suitable for use with these repeated measures data: (i) the summary statistic method; (ii) random/mixed effects modeling; (iii) GEE regression modeling; and (iv) survival analysis. CONCLUSIONS: Outcome studies in bipolar illness ideally should be longitudinal in orientation, obtain outcomes data frequently over extended times, and employ large study samples. Missing data problems can be expected, and data analytic methods must accommodate missingness.

Bipolar Disorder↗

Modern statistical methods for handling missing repeated measurements in obesity trial data: beyond LOCF.

This paper brings together some modern statistical methods to address the problem of missing data in obesity trials with repeated measurements. Such missing data occur when subjects miss one or more follow-up visits, or drop out early from an obesity trial. A common approach to dealing with missing data because of dropout is 'last observation carried forward' (LOCF). This method, although intuitively appealing, requires restrictive assumptions to produce valid statistical conclusions. We review the need for obesity trials, the assumptions that must be made regarding missing data in such trials, and some modern statistical methods for analysing data containing missing repeated measurements. These modern methods have fewer limitations and less restrictive assumptions than required for LOCF. Moreover, their recent introduction into current releases of statistical software and textbooks makes them more readily available to the applied data analyses.

Clinical Trials as Topic↗

Limitations of statistical measures of error in assessing the accuracy of continuous glucose sensors.

BACKGROUND: Various statistical methods are commonly used to assess the accuracy of near-continuous glucose sensors. The performance and reliability of these methods have not been well described. METHODS: We used computer simulation to describe the behavior of several statistical measures including error grid analysis, receiver operating characteristics, correlation, and repeated measures under varying conditions. Actual data from an inpatient accuracy study conducted by the Diabetes Research in Children Network (DirecNet) were also used to demonstrate these limitations. RESULTS: Sensors that were made artificially inaccurate by randomly shuffling the pairings to reference values still fell in Zone A or B 78% of the time for the Clarke grid and 79% of the time for the modified grid. Area under the curve values for these shuffled pairs averaged 64% for hypoglycemia and 68% for hyperglycemia. Continuous error grid analysis resulted in 75% of shuffled pairs designated as "Accurate Readings" or "Benign Errors." Correlation analysis gave inconsistent results for sensors simulated to have identical accuracies with values ranging from 0.50 to 0.96. Simplistic repeated-measures analyses accounting for subject effects, but ignoring temporal correlation patterns substantially inflated the probability of falsely obtaining a statistically significant result. In simulations where the null hypothesis was correct, 23% of observed P values were <0.05 and 12% of observed P values were <0.01. CONCLUSION: Commonly used statistical methods can give overly optimistic and/or inconsistent notions of sensor accuracy if results are not placed in proper context. Novel techniques are needed to assess the accuracy of near-continuous glucose sensors.

Adolescent↗

Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap.

MOTIVATION: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap. RESULTS: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices. AVAILABILITY: The sequence sets and code for performing these analyses are available from http://compbio.berkeley.edu/. CONTACT: brenner@compbio.berkeley.edu.

Algorithms↗

Tutorials in clinical research, part VI: descriptive statistics.

OBJECTIVES: This is the sixth in a series of Tutorials in Clinical Research. The objectives of this tutorial are to produce a brief, step-by-step, user-friendly atlas with clinical scenarios and illustrative examples to serve as ready references and, it is hoped, to open an inviting door to understanding statistics. STUDY DESIGN: Tutorial. METHODS: The authors met weekly for 6 months discussing clinical research articles and the applied statistics. Liberal use of reference texts and discussion focused on constructing a tutorial that is factual as well as easy to read and understand. RESULTS: The tutorial is organized into five sections: Overview, Key to Descriptive Summaries, Steps in Summarizing Descriptive Data, Understanding the Summaries of Specific Data Types, and Understanding Scope or Dispersion of Data. CONCLUSIONS: Assessing the validity of medical studies requires a working knowledge of statistics; however, obtaining this knowledge need not be beyond the ability of the busy surgeon. Desire is the only limiting factor. We have tried to construct an accurate, basic, easy-to-read-and-apply introduction to the task of summarizing the characteristics of a single variable. This is known as descriptive statistics.

Data Interpretation, Statistical↗

Common statistical methods in orthopaedic clinical studies.

We discuss the statistical representation and management of random error in orthopaedic clinical studies. Descriptive studies (such as case series) collect information about a sample that may be generalized to describe a population. Typically this description is in the form of summary statistics, such as means, proportions, or rates. Error in these variables may be represented by confidence intervals. Correlation and regression are techniques for investigation of the relationship between two or more variables. Descriptive statistics, comparisons of groups, especially hypothesis tests, and assessment of association including correlation and regression are important statistical concepts for clinicians involved in the conduct or appraisal of orthopaedic clinical research.

Chi-Square Distribution↗

Probability models and the applicability of statistical procedures in the identification of chromosomal fragile sites.

Böhm et al. (1995, Human Genetics 95, 249-256) introduced a statistical model (named FSM--fragile site model) specifically designed for the identification of fragile sites from chromosomal breakage data. In response to claims to the contrary (Hou et al., 1999, Human Genetics 104, 350-355; Hou et al., 2001, Biometrics 57, 435-440), we show how the FSM model is correctly modified for application under the assumption that the probability of random breakage is proportional to chromosomal band length and how the purportedly alternative procedures proposed by Hou, Chang, and Tai (1999, 2001) are variations of the correctly modified FSM algorithm. With the exception of the test statistic employed, the procedure described by Hou et al. (1999) is shown to be functionally identical to the correctly modified FSM and the application of an incorrectly modified FSM is shown to invalidate all of the comparisons of FSM to the alternatives proposed by Hou et al. (1999, 2001). Last, we discuss the statistical implications of the methodological variations proposed by Hou et al. (2001) and emphasize the logical and statistical necessity for fragile site identifications to be based on data from single individuals.

Algorithms↗

Perspectives on statistical significance testing.

The question of whether statistical significance testing should be used for the analysis of public health and epidemiologic data has received considerable attention in recent years. In this paper we have described some of the arguments for and against the use of hypothesis testing for the analysis of biomedical data. In addition, we have reviewed the literature from related fields, in particular sociology and psychology, in which similar discussions have taken place within the last 30 years. Many of the significance testing criticisms in these scientific fields have been raised in the more recent discussions taking place in the biomedical field. We present an example that emphasizes the use of both confidence interval estimation and significance testing. The example is particularly pertinent because it represents a more complex problem than has generally been discussed by critics of significance testing. Much of the discussion on this topic has focused on simple data analysis, such as the analysis of a 2 x 2 table or problems involving simple linear regression. Most epidemiologic data are far more complicated and warrant the use of both confidence interval estimation and significance testing for statistical analysis. Both of these techniques have no doubt been misused in the analysis of data. These misuses may have arisen from a lack of understanding of the role of statistical methods in data analysis and the choice of such methods for data analysis. If used prudently and judiciously, significance testing can help reduce the number of variables involved in a statistical analysis, thereby resulting in shorter confidence intervals for the models presented. Both significance testing and confidence interval estimation can serve and have served very useful functions for the analysis of public health and biomedical data.

Data Interpretation, Statistical↗

Between-subject and within-subject statistical information in dental research.

The evaluation of risk factors in dental research frequently uses observations at multiple sites in the same patient. For this reason, statistical methods that accommodate correlated data are generally used to assess the significance of the risk factors (e.g., generalized estimating equations, generalized linear mixed models). In applications of these methods, it is typically assumed (implicitly, if not explicitly) that between-subject and within-subject comparisons will produce the same estimated effect of the risk factor. When between- and within-subject comparisons conflict, the statistical methods can give biased estimates or results that are difficult to interpret. For illustration, we present two examples from periodontal disease studies in which different statistical methods give different estimates and significance levels for a risk factor. Statistical analyses in dental research should assess whether different sources of information give similar conclusions about risk factors or treatments.

Confounding Factors, Epidemiologic↗

Statistical issues in periodontal research.

In the past 10 to 12 years, there have been several statistical issues identified in periodontal research which require and have generated non-standard or new statistical approaches. The purpose of this paper is to give an overview of these issues and approaches. Three general categories of issues are described: (i) statistical methods for detecting when disease progression occurs, and biological theories and corresponding statistical models which attempt to describe how the disease progresses; (ii) design issues in studies of therapeutic efficacy; and (iii) analytic issues arising from periodontal data analysis.

Clinical Trials as Topic↗