Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

A primer for statistical analysis of clinical trials.

Randomized controlled trials have become the cornerstone of current practice evidence-based medicine. In this article, we present concise descriptions of fundamental statistical concepts and methods frequently used in the analysis of randomized trials. These include descriptive statistics, statistical inferences, techniques for the comparison of means or proportions from two samples, correlation, and regression analysis methods. The uses of these methods are illustrated with a number of practical examples, and the pitfalls of these topics are also briefly discussed. Lastly, some frequently used statistical terms and their meanings are also provided. By the end of the article, the reader should have sufficient knowledge to appreciate the statistical aspects of most clinical trial reports.

Evidence-Based Medicine↗

On the frontier of The Empire of Chance: statistics, accidents, and risk in industrializing America.

In The Empire of Chance, historians of science Gigerenzer et al. argue that statistical thinking has been "second to no other area of scientific endeavor" in its influence on "modern life and thought" (Gigerenzer et al. 1989, xiv-xv). This article describes how quantitative descriptions of risk associated with industrialization and technological change became part of the mentality of ordinary Americans. It explains why American began counting accidents, tells what kinds of accidents they counted and how they counted them, and shows how statistical representations of risk were used to justify prescriptions for public policy and individual behavior. On this frontier of the empire of chance, safety experts and self-styled "practical statisticians" were the principal colonizers. Distant from the centers of academic statistical science, they compromised rigorous scientific standards of methodology and accurate prediction in order to make convincing arguments outside their circles of expertise. To convey their point of view to audiences who were not literate in the field of statistics, they created a public language that conveyed statistical ideas through metaphors, graphic representations, and other rhetorical devices. They also engaged non-experts in collecting and analyzing data and, by the 1920s, even used quantitative self-measurement as a device to convince members of the public to alter their own risk-taking behaviors.

Accidents, Occupational↗

Statistical learning of higher-order temporal structure from visual shape sequences.

In 3 experiments, the authors investigated the ability of observers to extract the probabilities of successive shape co-occurrences during passive viewing. Participants became sensitive to several temporal-order statistics, both rapidly and with no overt task or explicit instructions. Sequences of shapes presented during familiarization were distinguished from novel sequences of familiar shapes, as well as from shape sequences that were seen during familiarization but less frequently than other shape sequences, demonstrating at least the extraction of joint probabilities of 2 consecutive shapes. When joint probabilities did not differ, another higher-order statistic (conditional probability) was automatically computed, thereby allowing participants to predict the temporal order of shapes. Results of a single-shape test documented that lower-order statistics were retained during the extraction of higher-order statistics. These results suggest that observers automatically extract multiple statistics of temporal events that are suitable for efficient associative learning of new temporal features.

Analysis of Variance↗

When clinical description becomes statistical prediction.

This article reconsiders the issue of clinical versus statistical prediction. The term clinical is widely used to denote 1 pole of 2 independent axes: the observer whose data are being aggregated (clinician/expert vs. lay) and the method of aggregating those data (impressionistic vs. statistical). Fifty years of research suggests that when formulas are available, statistical aggregation outperforms informal, subjective aggregation much of the time. However, these data have little bearing on the question of whether, or under what conditions, clinicians can make reliable and valid observations and inferences at a level of generality relevant to practice or useful as data to be aggregated statistically. An emerging body of research suggests that clinical observations, just like lay observations, can be quantified using standard psychometric procedures, so that clinical description becomes statistical prediction.

Cognition↗

Distant melodies: statistical learning of nonadjacent dependencies in tone sequences.

Human listeners can keep track of statistical regularities among temporally adjacent elements in both speech and musical streams. However, for speech streams, when statistical regularities occur among nonadjacent elements, only certain types of patterns are acquired. Here, using musical tone sequences, the authors investigate nonadjacent learning. When the elements were all similar in pitch range and timbre, learners acquired moderate regularities among adjacent tones but did not acquire highly consistent regularities among nonadjacent tones. However, when elements differed in pitch range or timbre, learners acquired statistical regularities among the similar, but temporally nonadjacent, elements. Finally, with a moderate grouping cue, both adjacent and nonadjacent statistics were learned, indicating that statistical learning is governed not only by temporal adjacency but also by Gestalt principles of similarity.

Auditory Perception↗

Score statistic to test for genetic correlation for proband-family design.

In genetic epidemiological studies informative families are often oversampled to increase the power of a study. For a proband-family design, where relatives of probands are sampled, we derive the score statistic to test for clustering of binary and quantitative traits within families due to genetic factors. The derived score statistic is robust to ascertainment scheme. We considered correlation due to unspecified genetic effects and/or due to sharing alleles identical by descent (IBD) at observed marker locations in a candidate region. A simulation study was carried out to study the distribution of the statistic under the null hypothesis in small data-sets. To illustrate the score statistic, data from 33 families with type 2 diabetes mellitus (DM2) were analyzed. In addition to the binary outcome DM2 we also analyzed the quantitative outcome, body mass index (BMI). For both traits familial aggregation was highly significant. For DM2, also including IBD sharing at marker D3S3681 as a cause of correlation gave an even more significant result, which suggests the presence of a trait gene linked to this marker. We conclude that for the proband-family design the score statistic is a powerful and robust tool for detecting clustering of outcomes.

Alleles↗

Statistical and deterministic approaches to designing transformations of electrocardiographic leads.

Two different approaches can be used to investigate the relationships among electrocardiographic leads: a statistical one, based on the analysis of recorded electrocardiograms (ECGs), and a deterministic one, based on physical principles that govern the current flow in irregularly shaped volume conductors such as the human body. The purpose of this study was to compare these two approaches. For the statistical investigation, the data set consisted of 120-lead ECGs recorded in a population including normal subjects (n = 290), post-myocardial-infarction patients (n = 497), patients with a history of ventricular tachycardia but no evidence of a previous myocardial infarction (n = 105), and patients with a single-vessel coronary artery disease who underwent coronary angioplasty (n = 91). Lead transformations of interest were obtained by fitting the multiple-regression model to this data set by the least-squares method. For the deterministic investigation, we used a boundary-element model of the human torso to simulate body-surface potentials in response to three orthogonal unit dipoles placed consecutively at 1,239 ventricular source locations, and the resulting body-surface potential distributions (instead of the recorded ECGs) were then fitted by the multiple-regression model. The results suggest that the lead transformations should be preferably designed by statistical analysis of recorded ECGs. Regression models with a small number of predictors (eg, those based on three ECG leads) are the most reliable; those using more predictors are fraught with the danger of collinearity when predictors are highly correlated (as occurs in the standard 12-lead ECG). Model-derived deterministic transformations are compatible with statistically derived ones, provided that the distributed character of the cardiac sources is taken into account. We conclude that statistical associations among electrocardiographic leads can be reliably quantified in sufficiently large and diverse databases of recorded data; the causality of these associations can be supported by appropriate deterministic models based on the laws of physics.

Angioplasty, Balloon, Coronary↗

[Use of statistical procedures in orthopedics exemplified by several years of an orthopedic specialty journal].

As expected, it was found that mainly graphic and tabular methods were used, unless the papers in question were purely descriptions of methods or brief case reports. This is probably connected with the long courses of orthopedic conditions, which, with a reasonable investment in time, can only be studied retrospectively. Nevertheless, it became apparent that the group of prospective studies, though numerically smaller, had by and large been performed quite well. Many of them managed without test statistics and yet have considerable information value. The few reports with test statistics had for the most part also been quite well conducted. The most common sources of errors in these were that too many tests were conducted on the same data material without taking the cumulative probability of error into account, e.g., according to Bonferroni-Holm; and the test method used was often not mentioned. So some good statistical studies in orthopedics are certainly being published, albeit gradually. It is planned to conduct a similar investigation on the same years' issues of a German-language journal in another specialty included in the list mentioned at the beginning--German to avoid the bias of translation. Afterwards the two studies can be compared to establish whether an orthopedic journal cannot also be included in the list of the 200. The necessity of impeccable statistics for practice and research is undisputed. The results presented here are intended to encourage orthopedists to attempt prospective studies more frequently than hitherto, and, keeping the test preconditions in mind, also to use correctly described, conclusive statistics.

Clinical Trials as Topic↗

Exploring the relationship between surrogates and clinical outcomes: analysis of individual patient data vs. meta-regression on group-level summary statistics.

There has been an increasing interest in exploring the relationship between a surrogate and a clinical outcome. Two different statistical approaches have been taken by researchers to quantify the treatment effect on the clinical outcome explained by the surrogate endpoint: 1) analysis based on individual patient data (IPD), and 2) meta-regression based on summary statistics from published literature. An analysis based on IPD models the associations between the surrogate and clinical outcome for patients directly and is able to adjust for patient-level covariates. A meta-regression models the trial-level associations using group-level summary statistics and trial-level covariates. The results from these two approaches can be quite disparate and researchers may reach different conclusions on scientific questions that they wish to answer. We demonstrate that the typical summary statistics, such as group means and event counts, do not provide a set of sufficient statistics for estimating the underlying relationship between the surrogate and clinical outcome for patients. Consequently, the associations derived from meta-regression do not necessarily reflect the causal relationship for patients and should be interpreted with caution. A meta-analysis of antiresorptive agents for osteoporosis serves to illustrate the magnitude of differences between the two approaches.

Humans↗

Monte Carlo dose calculations and radiobiological modelling: analysis of the effect of the statistical noise of the dose distribution on the probability of tumour control.

The aim of this work is to investigate the influence of the statistical fluctuations of Monte Carlo (MC) dose distributions on the dose volume histograms (DVHs) and radiobiological models, in particular the Poisson model for tumour control probability (tcp). The MC matrix is characterized by a mean dose in each scoring voxel, d, and a statistical error on the mean dose, sigma(d); whilst the quantities d and sigma(d) depend on many statistical and physical parameters, here we consider only their dependence on the phantom voxel size and the number of histories from the radiation source. Dose distributions from high-energy photon beams have been analysed. It has been found that the DVH broadens when increasing the statistical noise of the dose distribution, and the tcp calculation systematically underestimates the real tumour control value, defined here as the value of tumour control when the statistical error of the dose distribution tends to zero. When increasing the number of energy deposition events, either by increasing the voxel dimensions or increasing the number of histories from the source, the DVH broadening decreases and tcp converges to the 'correct' value. It is shown that the underestimation of the tcp due to the noise in the dose distribution depends on the degree of heterogeneity of the radiobiological parameters over the population; in particular this error decreases with increasing the biological heterogeneity, whereas it becomes significant in the hypothesis of a radiosensitivity assay for single patients, or for subgroups of patients. It has been found, for example, that when the voxel dimension is changed from a cube with sides of 0.5 cm to a cube with sides of 0.25 cm (with a fixed number of histories of 10(8) from the source), the systematic error in the tcp calculation is about 75% in the homogeneous hypothesis, and it decreases to a minimum value of about 15% in a case of high radiobiological heterogeneity. The possibility of using the error on the tcp to decide how many histories to run for a given voxel size is also discussed.

Computer Simulation↗

Use of runs statistics for pattern recognition in genomic DNA sequences.

In this article, the use of the finite Markov chain imbedding (FMCI) technique to study patterns in DNA under a hidden Markov model (HMM) is introduced. With a vision of studying multiple runs-related statistics simultaneously under an HMM through the FMCI technique, this work establishes an investigation of a bivariate runs statistic under a binary HMM for DNA pattern recognition. An FMCI-based recursive algorithm is derived and implemented for the determination of the exact distribution of this bivariate runs statistic under an independent identically distributed (IID) framework, a Markov chain (MC) framework, and a binary HMM framework. With this algorithm, we have studied the distributions of the bivariate runs statistic under different binary HMM parameter sets; probabilistic profiles of runs are created and shown to be useful for trapping HMM maximum likelihood estimates (MLEs). This MLE-trapping scheme offers good initial estimates to jump-start the expectation-maximization (EM) algorithm in HMM parameter estimation and helps prevent the EM estimates from landing on a local maximum or a saddle point. Applications of the bivariate runs statistic and the probabilistic profiles in conjunction with binary HMMs for pattern recognition in genomic DNA sequences are illustrated via case studies on DNA bendability signals using human DNA data.

Algorithms↗

Determinants of the autopsy decision: a statistical analysis.

Our goal was to use cross-sectional national mortality data to provide a multivariable statistical analysis of the factors that contribute to the decision of whether an autopsy will be performed. The identification of determinants of the autopsy is an important prerequisite for finding cost-effective alternatives for arresting or reversing the decline of autopsy rates in the circumstances in which the autopsy can continue to make a crucial contribution to clinical medicine and public health. The source of the data was 1986 National Center for Health Statistics (Washington, DC) mortality data tapes for Kentucky, Maryland, Minnesota, and Washington for the 1986 calendar year. Separate multiple logistic regressions were conducted on these data on a state-by-state basis, with a total of 139,063 individual mortality records as the unit of analysis. The dependent variable in all models was autopsy (yes/no). Odds ratios for selected explanatory variables were estimated for all four states, and the relative contribution of each explanatory variable was studied in a detailed analysis of one state. In general, the following independent variables had a statistically significant positive relationship with whether an autopsy will be performed: male sex; nonwhite ethnicity; death due to ill-defined or unknown cause; death due to accident, suicide, or homicide; presence of a nationally recognized medical center in the county of death; and death occurring in a standard metropolitan statistical area. In general, the following independent variables had a statistically significant negative relationship with whether an autopsy will be performed: older age at death; higher income level of the decedent; death in a nursing home; death at home; and residency in the county of death. The two most important variables influencing the autopsy decision were age at death (especially old age) and death due to accident, homicide, or suicide.

Adolescent↗

SpA: web-accessible spectratype analysis: data management, statistical analysis and visualization.

SUMMARY: SpA is a web-accessible system for the management, visualization and statistical analysis of T-cell receptor spectratype data. Users upload data from their spectratype analyzers to SpA, which saves the raw data and user-defined supplementary covariates to a secure database. The statistical engine performs several data analyses and statistical summaries. The visualization engine displays spectratype histograms in a Java applet and in an image file suitable for download. All of these results are also saved to the database and remain accessible to the user. Additional statistical tools specific to the analysis of multiple spectratypes are also available through the SpA interface. AVAILABILITY: The service is freely accessible via the web at http://www.duke.edu/~kepler/spa.html. Additional technical support and specialized statistical analysis and consultation are available by arrangement with the authors and, depending on the service requested, may be subject to fee.

Animals↗

Responding to community-identified suicide clusters: statistical verification of the cluster is not the primary issue.

Establishing the presence of an epidemic is traditionally a first step in any outbreak investigation. For two reasons, however, this has not been a fruitful approach for suicide cluster investigations. First, the data necessary to statistically verify an excess number of suicidal incidents are often lacking or of poor quality. Second, and more important, when a community perceives that it is experiencing a suicide cluster, it is not immediately relevant whether the cluster is statistically significant. The perception of suicide clustering, and the highly charged emotional atmosphere associated with that perception, may dramatically heighten the potentially "contagious" effect of suicide. That the perception of clustering may itself be a risk factor for suicide distinguishes suicide clusters from all other clusters of fatal disease or illness. A community response plan should, therefore, be implemented to identify and refer persons who may be at high risk of suicide, regardless of whether the community-identified suicide cluster is statistically significant. Statistical techniques may be useful at several stages in the investigation and control of apparent suicide clusters, but statistical verification of a community-identified suicide cluster is not appropriate as a starting point for response to the cluster.

Adolescent↗

Statistical evaluation of ventilator-free days as an efficacy measure in clinical trials of treatments for acute respiratory distress syndrome.

OBJECTIVE: Trials of potential new therapies in acute lung injury are difficult and expensive to conduct. This article is designed to determine the utility, behavior, and statistical properties of a new primary end point for such trials, ventilator-free days, defined as days alive and free from mechanical ventilation. Describing the nuances of this outcome measure is particularly important because using it, while ignoring mortality, could result in misleading conclusions. DESIGN: To develop a model for the duration of ventilation and mortality and fit the model by using data from a recently completed clinical trial. To determine the appropriate test statistic for the new measure and derive a formula for power. To determine a formula for the probability that the test statistic will reject the null hypothesis and mortality will simultaneously show improvement. To plot power curves for the test statistic and determine sample sizes for reasonable alternative hypotheses. SETTING: Intensive care units. PATIENTS: Patients with acute respiratory distress syndrome or acute lung injury as defined by the American-European Consensus Conference. MAIN RESULTS: The proposed model fit the clinical data. Ventilator-free days were improved by lower tidal volume ventilation, but the improvement was mostly caused by the improved mortality rate, so trials that expected similar effects would only have modest increase in power if they used ventilator-free days as their primary end point rather than 28-day mortality. Similar results were obtained using the model in two groups segregated by low or high Acute Physiology and Chronic Health Evaluation score. On the other hand, if patients are divided into two groups on the basis of the lung injury score, both the duration of ventilation and mortality are lower in the low lung injury score group. A trial of a treatment that had a similar clinical effect would have a large increase in power, allowing for a reduction in the required sample size. CONCLUSIONS: Use of ventilator-free days as a trial end point allows smaller sample sizes if it is assumed that the treatment being tested simultaneously reduces the duration of ventilation and improves mortality. It is unlikely that a treatment that led to higher mortality could lead to a statistically significant improvement in ventilator-free days. This would be especially true if the treatment were also required to produce a nominal improvement in mortality.

APACHE↗

Association analysis of polymorphisms in serotonin 1B receptor (HTR1B) gene with heroin addiction: a comparison of molecular and statistically estimated haplotypes.

OBJECTIVES: 5-Hydroxytryptamine (serotonin)-1B receptors (HTR1B) may play an important role in psychiatric disorders and drug and alcohol dependence. In this study we report on genotype, molecular haplotype and statistically estimated haplotype analyses of previously identified polymorphisms in positions -261T>G, -161A>T, 129C>T, 861G>C and 1180A>G of the HTR1B gene in ethnically diverse populations (African-Americans, Caucasians, Hispanics and Asians) including 235 former heroin addicts and 161 control subjects from New York City. The objectives were to test for an association of molecular and statistically estimated haplotypes and genotypes in HTR1B gene with heroin addiction and to compare results provided by molecular and statistically estimated haplotyping methods. METHODS: Genotype analysis was performed using a standard TaqMan protocol. Molecular haplotype analysis of the subset of polymorphisms consisting of -261T>G, -161A>T and 129C>T was performed using a protocol specially designed by our group, using fluorescent PCR. This is based on use of allele-specific primers complementary to flanking polymorphisms and a fluorescently labeled sequence-specific TaqMan probe set complementary to an internal polymorphism of the haplotype region. Every individual's statistically inferred haplotype pair agreed with the individual's haplotype pair determined by molecular haplotyping. RESULTS AND CONCLUSION: A point-wise significant association of haplotype pairs containing allele G at position 1180 with protective effect from heroin addiction in Caucasians was found. A point-wise nominally significant association of allele 1180G with a protective effect from heroin addiction was found in Caucasians. Statistically significant differences across four ethnic groups in control subjects for allelic frequencies of -261T>G and -161A>T were found.

Black or African American↗

Features of statistical dynamics in a finite system.

We study features of statistical dynamics in a finite Hamilton system composed of a relevant one degree of freedom coupled to an irrelevant multidegree of freedom system through a weak interaction. Special attention is paid on how the statistical dynamics changes depending on the number of degrees of freedom in the irrelevant system. It is found that the macrolevel statistical aspects are strongly related to an appearance of the microlevel chaotic motion, and a dissipation of the relevant motion is realized passing through three distinct stages: dephasing, statistical relaxation, and equilibrium regimes. It is clarified that the dynamical description and the conventional transport approach provide us with almost the same macrolevel and microlevel mechanisms only for the system with a very large number of irrelevant degrees of freedom. It is also shown that the statistical relaxation in the finite system is an anomalous diffusion and the fluctuation effects have a finite correlation time.

Journal Article↗

Classical singularities and semi-Poisson statistics in disordered systems.

We investigate a one-dimensional disordered Hamiltonian with a nonanalytical dispersion relation whose level statistics is exactly described by semi-Poisson statistics. It is shown that this result is robust, namely, it does not depend on the microscopic details of the Hamiltonian but only on the type of nonanalytical potential. We also argue that a deterministic kicked rotator with a steplike potential has the same spectral properties. Semi-Poisson statistics, typical of pseudointegrable billiards, have been frequently claimed to describe critical statistics, namely, the level statistics of a disordered system at the Anderson transition. However, we provide convincing evidence they are indeed different: each of them has its origin in a different type of classical singularity.

Journal Article↗