Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “STATISTICS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗

Linearly divergent treatment effects in clinical trials with repeated measures: efficient analysis using summary statistics.

In many randomized clinical trials with repeated measures of a response variable one anticipates a linear divergence over time in the difference between treatments. This paper explores how to make an efficient choice of analysis based on individual patient summary statistics. With the objective of estimating the mean rate of treatment divergence the simplest choice of summary statistic is the regression coefficient of response on time for each subject (SLOPE). The gains in statistical efficiency imposed by adjusting for the observed pre-treatment levels, or even better the estimated intercepts, are clarified. In the process, we develop the optimal linear summary statistic for any repeated measures design with assumed known covariance structure and shape of true mean treatment difference over time. Statistical power considerations are explored and an example from an asthma trial is used to illustrate the main points.

Analysis of Variance↗

Statistical shortcomings in licensing applications.

This paper concerns the statistical work carried out with respect to clinical trials conducted for regulatory purposes. Although the general quality of such work has improved markedly over recent years and is now generally high, a number of shortcomings remain. A few of these arise from failure to follow well established statistical practice. Rather more arise from a poor understanding of areas of known statistical disagreement and from the unsatisfactory use of newer and more advanced techniques. Inadequacies in reporting statistical work are commonplace. Examples of all these shortcomings are provided and emphasis is placed on the value of a statistical contribution to overall summaries such as the clinical expert report.

Clinical Trials, Phase II as Topic↗

Visit-driven endpoints in randomized HIV/AIDS clinical trials: impact of missing data on treatment difference measured on summary statistics.

In randomized HIV/AIDS clinical trials, CD4 lymphocyte counts and plasma HIV-1 RNA measurements are often used as endpoints. The comparison between treatment groups is mainly based on a summary measure of outcome, so-called summary statistic. Such analyses are often complicated by missing data occurring as drop-outs. For the most currently used summary statistics in these trials, we examined the impact of missing data occurring as drop-outs on test size, in order to help choosing between these statistics. A simulation of missing-data patterns was performed, using HIV-1 plasma RNA measurements as the main endpoint, to compare the effect of three plausible informative patterns, depending on treatment group, and on baseline or current plasma viral load, on eight different summary statistics. Missing data resulted in test sizes over the nominal value for the area under the curve minus baseline, the least-squares slope, the slope estimated with use of a mixed effects linear model, assuming a linear trend over the entire study, the difference between baseline and nadir, and the difference between baseline and week 24. The difference between baseline and week 8 was an acceptable summary with respect to the test size, but did not reflect accurately the durability of the effect of treatment. Two criteria appeared as the best summary statistics: the slope estimated by a mixed effects model, with a change of slope after two weeks of treatment, and to a lesser degree, the area under the curve after carrying forward the last observation.

Acquired Immunodeficiency Syndrome↗

Two-stage global search designs for linkage analysis I: use of the mean statistic for affected sib pairs.

Two-stage global search designs for linkage analysis using pairs of affected relatives were shown by Elston et al. [1996] to typically halve the cost of a study compared to a one-stage design. The statistic used for testing linkage in that study was based on the proportion of pairs sharing no marker alleles identical by descent (IBD). However, it has been established that the mean statistic often has the greatest power for full sib pairs [Blackwelder and Elston, 1985; Schaid and Nick, 1990; Knapp et al. 1994]. In this paper, we study optimal two-stage global search designs, in the case of affected full sib pairs, when using the mean test statistic to test for linkage. When dominant genetic variance is present, using the mean statistic is usually more cost efficient than using the proportion of pairs sharing no maker alleles IBD; in the case when there is no dominant genetic variance, the mean statistic leads to a better design, in the sense of being more cost-saving, provided that the relative risk ratio for first-degree relatives is small. The effect of heterogeneity and markers' informativeness is also investigated, the latter using the Linkage Information Content value for sibs.

Alleles↗

At what price significance? The effect of price estimates on statistical inference in economic evaluation.

Because data on resource utilization are now collected in many comparative trials of health interventions, statistical analysis of between-group differences in mean costs has become common. Statistical analyses of costs are generally performed conditional on a set of resource prices (or unit costs), thereby suppressing any uncertainty associated with those price estimates. Results presented here demonstrate that varying price estimates can have a non-negligible effect on statistical inference regarding between-group cost differences. Depending on the relative prices used in an analysis, between-group differences in total costs per patient may be either statistically significant or insignificant, regardless of whether differences in utilization of the underlying resources are statistically significant. These results highlight the importance of recognizing that evaluations based on patient-level economic data may be sensitive to assumptions regarding the values of unobserved variables, such as the relative prices of resources. Traditional methods of sensitivity analysis remain a valuable tool for analysing the implications of uncertainty around estimates of those unobserved variables.

Clinical Trials as Topic↗

Human performance and physiology: a statistical power analysis of ELF electromagnetic field research.

Research examining the effects of electromagnetic fields (EMFs) on human performance and physiology has produced inconsistent results; this might be attributable to low statistical power. Statistical power refers to the probability of obtaining a statistically significant result, given the fact that a real effect exists. The results of a survey of published investigations of the effects of EMFs on human performance and physiology show that statistical power levels are very low, ranging from a mean of .08 for small effect sizes to .46 for large effect sizes. Implications of these findings for the interpretation of results are discussed along with suggestions for increasing statistical power.

Behavior↗

Influences on inferences. Effect of errors in data on statistical evaluation.

BACKGROUND: Inadvertent random and systemic errors introduced into data sets and manipulation of data are well-defined sources of discrepancies in statistical evaluation of clinical trials. In this study, the authors show the influence of errors on the widely used statistical result, P values. METHODS: Using data from a retrospective study of patients with Hodgkin disease treated at the University of Minnesota between 1970 and 1984 and observed to 1988, we introduced various errors into the data to study the impact on results. RESULTS: Inadvertent random and systemic errors affect statistical results. Data entry and transcription errors, vague definitions of endpoints and prognostic factors, and the omission and selection of patients are examples of frequent errors that affect statistical evaluation. CONCLUSION: The results and inferences of many studies are sensitive to systemic errors and data manipulation. Great care must be given to the clear definitions of terms, exclusion and inclusion criteria, group assignments, treatment protocols, and the subgroups on which statistical analysis is performed. Clinicians and statisticians must work together to improve the performance and interpretation of clinical trials.

Clinical Trials as Topic↗

Misuse of statistical methods in Arthritis and Rheumatism. 1982 versus 1967-68.

Articles published in Arthritis and Rheumatism in 1982 were compared with those from 1967-68 to evaluate trends in statistical methods and in the quantity and character of statistical misuse. Results show that among articles in 1982 using statistics, 66% contained methodologic errors. The percentage was similar in 1967-68, although fewer articles used statistics. In 1967-68 the most common error was failure to identify the statistical method used, whereas in 1982 multiple testing errors predominated, specifically, use of the t-test to compare 3 or more groups and comparison of 2 groups on multiple variables. The availability of calculators and computers which readily perform complex data analysis may underly the emergence of multiple testing errors. To compare multiple groups, we suggest using analysis of variance instead of the t-test. To avoid multiple testing errors, we recommend limiting the number of tests performed, lowering the alpha (significance) level, or using multivariate techniques.

Arthritis↗

An empirical analysis of eating disorders and anxiety disorders publications (1980-2000)--part II: Statistical hypothesis testing.

OBJECTIVE: The current study compared the eating disorder literature and the anxiety disorder literature in terms of statistical hypothesis testing features in 1980, 1990, and 2000. METHOD: Computer literature searches were conducted using PubMed and PsychInfo databases to identify relevant eating disorder and anxiety disorder articles published at each of the three time points. A total of 456 articles were randomly selected, including 228 articles each from the fields of eating disorders and anxiety disorders. Within each field, one third (76) of the articles were selected from each of the three time points. Two raters, from a team of eight trained raters, were randomly assigned to independently rate each article in terms of 75 separate methodologic features. In the current article, we will emphasize the findings about hypothesis testing and statistical analysis. Disagreements in ratings were resolved via consensus. Ratings were tabulated separately by field across the three time points. RESULTS: Few differences were observed between eating disorder and anxiety disorder publications in terms of statistical hypothesis testing features. Although increases were observed in both fields in a number of areas from 1980 to 2000, there remains a pervasive absence of many of the statistical hypothesis testing features recommended by the American Psychological Association Task Force on Statistical Inference. CONCLUSION: These results are discussed in terms of their implications for the fields of eating disorders and anxiety disorders, for researchers, for reviewers, and for professional journals and editorial boards.

Anxiety Disorders↗

Statistical tests of significance in transgenic mutation assays: considerations on the experimental unit.

When significant animal-to-animal variability is present in binary response data, the usual statistical tests applied to such data do not always operate correctly. In transgenic mouse mutation data, some evidence of significant animal-to-animal variability already exists, suggesting that conventional statistical methods may not be appropriate. Here, we describe an alternative statistical method that treats the animal as the experimental (or statistically independent) unit, and contrast results of its application with those from methods that take the transgene as the experimental unit. Using data from two publications that report experimental results for individual animals, the transgene-based and animal-based analyses can yield very different interpretations of the experimental data. The performance of animal-based statistical methods should be improved by conducting future experiments with enough animals to adequately address animal-to-animal variability.

Animals↗

A Monte Carlo evaluation of three statistical methods used in path analysis.

Results of a Monte Carlo study to investigate the properties of three statistical methods used extensively in path analysis of family data are presented. All three methods are based on the maximum likelihood principle and involve the assumptions of multivariate normality and large sample (asymptotic) statistical properties. The methods differ, however, in the specification of the likelihood function. Given a set of correlation estimates, method 1 maximizes the likelihood function under the stipulation that the estimates are independent. Method 2 differs from the former by allowing for covariances among the correlation estimators. Method 3 involves (direct) maximization of the likelihood function for the individual family observations assuming multivariate normality for the vector of family observations. The Monte Carlo study investigated validity of the test statistics and confidence intervals and evaluated the relative efficiency and bias of the parameter estimates based on 1,000 replications of each of several simulation conditions. The effects of violating the two basic assumptions, multivariate normality and asymptotic theory, were investigated by comparing results for non-normally vs normally distributed family data and for small vs large sample sizes. It is shown that method 3 provides valid statistical inferences under multivariate normality and that it is generally robust against minor departures from normality. Method 2 is also robust against minor deviations from normality, but it is sensitive to small sample sizes. Method 1 yields highly conservative test statistics under all conditions studied.

Genetics, Medical↗

The weighted rank pairwise correlation statistic for linkage analysis: simulation study and application to Alzheimer's disease.

The weighted rank pairwise correlation (WRPC) statistic has been proposed as a robust test of genetic linkage, particularly adapted to the analysis of large and complex pedigrees and for age-dependent and heterogeneous diseases. In this paper a simulation study is presented. Validity and power of the WRPC test are studied and compared to the Haseman-Elston sibpair method for various types of problems. The power of the WRPC test is slightly lower than the Haseman-Elston method for analyzing a large number of small randomly chosen pedigrees. It is higher however in presence of genetic heterogeneity or for analyzing large individual pedigrees. Recently, evidence of linkage of Alzheimer's disease with a locus on chromosome 14, D14s43, has been obtained by the Lod-score method. We reanalyze these data using the WRPC test, essentially confirming the results of the Lod-score method. The WRPC test statistic is higher than the equivalent Lod-score statistic for the two pedigrees which show strong evidence of linkage with the two methods. The global WRPC test statistic is slightly lower than the Lod-score test statistic. The WRPC test, however, makes no hypothesis of a specific genetic transmission model and can be computed very quickly; in addition, an exact P-value can be computed by simulation for individual pedigrees.

Alzheimer Disease↗

Clinical versus statistical prediction: the contribution of Paul E. Meehl.

The background of Paul E. Meehl's work on clinical versus statistical prediction is reviewed, with detailed analyses of his arguments. Meehl's four main contributions were the following: (a) he put the question, of whether clinical or statistical combinations of psychological data yielded better predictions, at center stage in applied psychology; (b) he convincingly argued, against an array of objections, that clinical versus statistical prediction was a real (not concocted) problem needing thorough study; (c) he meticulously and even-handedly dissected the logic of clinical inference from theoretical and probabilistic standpoints; and (c) he reviewed the studies available in 1954 and thereafter, which tested the validity of clinical versus statistical predictions. His early conclusion that the literature strongly favors statistical prediction has stood up extremely well, and his conceptual analyses of the prediction problem (especially his defense of applying aggregate-based probability statements to individual cases) have not been significantly improved since 1954.

Forecasting↗

The Statistical Fragility of Saline Nasal Irrigation for Rhinosinusitis: A Systematic Review.

OBJECTIVE: To assess the statistical fragility of randomized controlled trials (RCTs) evaluating high-volume saline nasal irrigation (SNI) for rhinosinusitis using fragility analysis. DATA SOURCES: PubMed, MEDLINE, and Embase were searched for RCTs published between May 1976 and January 2026. REVIEW METHODS: This study was reported as per PRISMA guidelines. RCTs that compared high-volume SNI to non-irrigation standard care for acute, recurrent, or chronic rhinosinusitis, and reported ≥ 1 dichotomous outcome, were included. Fragility index (FI), the minimum number of event reversals needed to alter statistical significance, and fragility quotient (FQ), FI normalized to sample size, were calculated for statistically significant dichotomous outcomes. Reverse FI (rFI) and reverse FQ (rFQ) were calculated for non-significant outcomes. RESULTS: Eight RCTs were included, yielding 38 dichotomous outcomes. Eight outcomes (21.1%) were statistically significant. The overall combined median FI was 5 (FQ 0.062), with similar FI values between significant and non-significant outcomes. In over one-fifth of outcomes, loss to follow-up exceeded FI. Analysis of principal dichotomous outcomes from studies demonstrated a median FI of 6 (FQ 0.092), with five of eight (62.5%) outcomes non-significant. CONCLUSION: RCTs evaluating SNI for rhinosinusitis exhibit moderate-to-high statistical fragility, with small outcome changes capable of reversing study conclusions. Because fragility analysis was limited to dichotomous outcomes while many primary endpoints were continuous, our findings should be interpreted as complementary rather than comprehensive appraisals of RCTs. Future RCTs with larger sample sizes, reduced bias, and pre-specified fragility considerations are needed to better define the clinical role of SNI.

Rhinosinusitis↗

Normal mode analysis of macromolecular motions in a database framework: developing mode concentration as a useful classifying statistic.

We investigated protein motions using normal modes within a database framework, determining on a large sample the degree to which normal modes anticipate the direction of the observed motion and were useful for motions classification. As a starting point for our analysis, we identified a large number of examples of protein flexibility from a comprehensive set of structural alignments of the proteins in the PDB. Each example consisted of a pair of proteins that were considerably different in structure given their sequence similarity. On each pair, we performed geometric comparisons and adiabatic-mapping interpolations in a high-throughput pipeline, arriving at a final list of 3,814 putative motions and standardized statistics for each. We then computed the normal modes of each motion in this list, determining the linear combination of modes that best approximated the direction of the observed motion. We integrated our new motions and normal mode calculations in the Macromolecular Motions Database, through a new ranking interface at http://molmovdb.org. Based on the normal mode calculations and the interpolations, we identified a new statistic, mode concentration, related to the mathematical concept of information content, which describes the degree to which the direction of the observed motion can be summarized by a few modes. Using this statistic, we were able to determine the fraction of the 3,814 motions where one could anticipate the direction of the actual motion from only a few modes. We also investigated mode concentration in comparison to related statistics on combinations of normal modes and correlated it with quantities characterizing protein flexibility (e.g., maximum backbone displacement or number of mobile atoms). Finally, we evaluated the ability of mode concentration to automatically classify motions into a variety of simple categories (e.g., whether or not they are "fragment-like"), in comparison to motion statistics. This involved the application of decision trees and feature selection (particular machine-learning techniques) to training and testing sets derived from merging the "list" of motions with manually classified ones.

Databases, Protein↗

A workflow spatial scan statistic.

We propose a modification of the spatial scan statistic that takes account of workflow, which is the movement of individuals between home and work. The objective is to detect clusters of disease in situations where exposure occurs in the workplace, but only home address is available for analysis. In these situations, application of the usual spatial scan statistic does not account for possible differences between home and work address, thereby reducing the power of detection. We describe an extension to the usual spatial scan statistic that uses workflow data to search for disease clusters resulting from workplace exposure. We also present results from simulations that demonstrate the increased power of the workflow scan statistic over the usual scan statistic for detecting clusters arising from exposures in the workplace.

Anthrax↗

The geography of power: statistical performance of tests of clusters and clustering in heterogeneous populations.

Heterogeneous population densities complicate comparisons of statistical power between hypothesis tests evaluating spatial clusters or clustering of disease. Specifically, the location of a cluster within a heterogeneously distributed population at risk impacts power properties, complicating comparisons of tests, and allowing one to map spatial variations in statistical power for different tests. Such maps provide insight into the overall power of a particular test, and also indicate areas within the study area where tests are more or less likely to detect the same local increase in relative risk. While such maps are largely driven by local sample size, we also find differences due to features of the statistics themselves. We illustrate these concepts using two tests: Tango's index of clustering and the spatial scan statistic. Furthermore, assessments of the accuracy of the 'most likely cluster' involve not only statistical power, but also spatial accuracy in identifying the location of a true underlying cluster. We illustrate these concepts via induction of artificial clusters within the observed incidence of severe cardiac birth defects in Santa Clara County, CA in 1981.

California↗