Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “statistical inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Précis of statistical significance: rationale, validity, and utility.

The null-hypothesis significance-test procedure (NHSTP) is defended in the context of the theory-corroboration experiment, as well as the following contrasts: (a) substantive hypotheses versus statistical hypotheses, (b) theory corroboration versus statistical hypothesis testing, (c) theoretical inference versus statistical decision, (d) experiments versus nonexperimental studies, and (e) theory corroboration versus treatment assessment. The null hypothesis can be true because it is the hypothesis that errors are randomly distributed in data. Moreover, the null hypothesis is never used as a categorical proposition. Statistical significance means only that chance influences can be excluded as an explanation of data; it does not identify the nonchance factor responsible. The experimental conclusion is drawn with the inductive principle underlying the experimental design. A chain of deductive arguments gives rise to the theoretical conclusion via the experimental conclusion. The anomalous relationship between statistical significance and the effect size often used to criticize NHSTP is more apparent than real. The absolute size of the effect is not an index of evidential support for the substantive hypothesis. Nor is the effect size, by itself, informative as to the practical importance of the research result. Being a conditional probability, statistical power cannot be the a priori probability of statistical significance. The validity of statistical power is debatable because statistical significance is determined with a single sampling distribution of the test statistic based on H0, whereas it takes two distributions to represent statistical power or effect size. Sample size should not be determined in the mechanical manner envisaged in power analysis. It is inappropriate to criticize NHSTP for nonstatistical reasons. At the same time, neither effect size, nor confidence interval estimate, nor posterior probability can be used to exclude chance as an explanation of data. Neither can any of them fulfill the nonstatistical functions expected of them by critics.

Reproducibility of Results↗

Single-trial lambda wave identification using a fuzzy inference system and predictive statistical diagnosis.

The aim of the study was to automate the identification of a saccade-related visual evoked potential (EP) called the lambda wave. The lambda waves were extracted from single trials of electroencephalogram (EEG) waveforms using independent component analysis (ICA). A trial was a set of EEG waveforms recorded from 64 scalp electrode locations while a saccade was performed. Forty saccade-related EEG trials (recorded from four normal subjects) were used in the study. The number of waveforms per trial was reduced from 64 to 22 by pre-processing. The application of ICA to the resulting waveforms produced 880 components (i.e. 4 subjects x 10 trials per subject x 22 components per trial). The components were divided into 373 lambda and 507 nonlambda waves by visual inspection and then they were represented by one spatial and two temporal features. The classification performance of a Bayesian approach called predictive statistical diagnosis (PSD) was compared with that of a fuzzy logic approach called a fuzzy inference system (FIS). The outputs from the two classification approaches were then combined and the resulting discrimination accuracy was evaluated. For each approach, half the data from the lambda and nonlambda wave categories were used to determine the operating parameters of the classification schemes while the rest (i.e. the validation set) were used to evaluate their classification accuracies. The sensitivity and specificity values when the classification approaches were applied to the lambda wave validation data set were as follows: for the PSD 92.51% and 91.73% respectively, for the FIS 95.72% and 89.76% respectively, and for the combined FIS and PSD approach 97.33% and 97.24% respectively (classification threshold was 0.5). The devised signal processing techniques together with the classification approaches provided for an effective extraction and classification of the single-trial lambda waves. However, as only four subjects were included, it will be valuable to further evaluate the methods on a larger group of subjects.

Adult↗

Blomqvist revisited: how and when to test the relationship between level and longitudinal rate of change.

Longitudinal studies are often interested in assessing the relationship between severity (level) and rate of change (slope). Blomqvist describes an estimator of this relationship that has been used in a variety of contexts. This paper reviews and generalizes the Blomqvist method. Most published applications of the Blomqvist method contain substantial bias because they fail to consider and accommodate confounding due to the pooling of multiple age cohorts in a single analysis. We describe this bias, and present an unbiased algorithm consistent with the intentions of Blomqvist. We also explore when it is appropriate to apply the Blomqvist analysis, and what inferences can be made using this statistic. Aetiological inference about premorbid level of function predicting future rate of decline is often desired, but may not be justified when modelling chronic progressive conditions, since differential progression prior to the start of longitudinal follow-up can induce a relationship between level and rate of decline, even in the absence of an aetiologically relevant association. We conclude that aetiological inference by the Blomqvist analysis is not appropriate in most investigations of chronic progressive disease. Using the model to develop descriptive and predictive equations in these circumstances, however, remains appropriate, as does testing simply for clinical heterogeneity in longitudinal rate of decline.

Aged↗

[The meaning of statistical data in medical science and their examination--true and false analysis of statistical data].

The subjects which are often encountered in the statistical design and analysis of data in medical science studies were discussed. The five topics examined were: Medical science and statistical methods So-called mathematical statistics and medical science Fundamentals of cross-tabulation analysis of statistical data and inference Exploratory study by multidimensional data analyses Optimal process control of individual, medical science and informatics of statistical data In I, the author's statistico-mathematical idea is characterized as the analysis of phenomena by statistical data. This is closely related to the logic, methodology and philosophy of science. This statistical concept and method are based on operational and pragmatic ideas. Self-examination of mathematical statistics is particularly focused in II and III. In II, the effectiveness of experimental design and statistical testing is thoroughly examined with regard to the study of medical science, and the limitation of its application is discussed. In III the apparent paradox of analysis of cross-tabulation of statistical data and statistical inference is shown. This is due to the operation of a simple two- or three-fold cross-tabulation analysis of (more than two or three) multidimensional data, apart from the sophisticated statistical test theory of association. In IV, the necessity of informatics of multidimensional data analysis in medical science is stressed. In V, the following point is discussed. The essential point of clinical trials is that they are not based on any simple statistical test in a traditional experimental design but on the optimal process control of individuals in the information space of the body and mind, which is based on a knowledge of medical science and the informatics of multidimensional statistical data analysis.

Statistics as Topic↗

On assessing interrater agreement for multiple attribute responses.

New methods are developed for assessing the extent of interrater agreement when each unit to be rated is characterized by a (possibly empty) subset of a specified set of distinct nominal attributes. For such multiple attribute response data, a two-rater concordance statistic is derived, and associated statistical inference-making procedures are provided. This concordance statistic is corrected for chance agreement by using an underlying hypergeometric model. Numerical examples are given to illustrate the proposed methodology, and comparisons to other agreement statistics (e.g., kappa) are made.

Biometry↗

Statistical power and parameter stability when subjects are few and tests are many: comment on Peterson, Smith, Martorana, and Owens (2003).

Comments on the original article "The impact of chief executive officer personality on top management team dynamics: One mechanism by which leadership affects organizational performance", by R. S. Peterson et al.. This comment illustrates how small sample sizes, when combined with many statistical tests, can generate unstable parameter estimates and invalid inferences. Although statistical power for 1 test in a small-sample context is too low, the experimentwise power is often high when many tests are conducted, thus leading to Type I errors that will not replicate when retested. This comment's results show how radically the specific conclusions and inferences in R. S. Peterson, D. B. Smith, P. V. Martorana, and P. D. Owens's (2003) study changed with the inclusion or exclusion of 1 data point. When a more appropriate experimentwise statistical test was applied, the instability in the inferences was eliminated, but all the inferences become nonsignificant, thus changing the positive conclusions.

Data Interpretation, Statistical↗

A statistical framework for haplotype block inference.

The existence of haplotype blocks transmitted from parents to offspring has been suggested recently. This has created an interest in the inference of the block structure and length. The motivation is that haplotype blocks that are characterized well will make it relatively easier to quickly map all the genes carrying human diseases. To study the inference of haplotype block systematically, we propose a statistical framework. In this framework, the optimal haplotype block partitioning is formulated as the problem of statistical model selection; missing data can be handled in a standard statistical way; population strata can be implemented; block structure inference/hypothesis testing can be performed; prior knowledge, if present, can be incorporated to perform a Bayesian inference. The algorithm is linear in the number of loci, instead of NP-hard for many such algorithms. We illustrate the applications of our method to both simulated and real data sets.

Algorithms↗

Valid conjunction inference with the minimum statistic.

In logic a conjunction is defined as an AND between truth statements. In neuroimaging, investigators may look for brain areas activated by task A AND by task B, or a conjunction of tasks (Price, C.J., Friston, K.J., 1997. Cognitive conjunction: a new approach to brain activation experiments. NeuroImage 5, 261-270). Friston et al. (Friston, K., Holmes, A., Price, C., Buchel, C., Worsley, K., 1999. Multisubject fMRI studies and conjunction analyses. NeuroImage 10, 85-396) introduced a minimum statistic test for conjunction. We refer to this method as the minimum statistic compared to the global null (MS/GN). The MS/GN is implemented in SPM2 and SPM99 software, and has been widely used as a test of conjunction. However, we assert that it does not have the correct null hypothesis for a test of logical AND, and further, this has led to confusion in the neuroimaging community. In this paper, we define a conjunction and explain the problem with the MS/GN test as a conjunction method. We present a survey of recent practice in neuroimaging which reveals that the MS/GN test is very often misinterpreted as evidence of a logical AND. We show that a correct test for a logical AND requires that all the comparisons in the conjunction are individually significant. This result holds even if the comparisons are not independent. We suggest that the revised test proposed here is the appropriate means for conjunction inference in neuroimaging.

Arousal↗

Statistical design, analysis, and inference issues in studies using sister chromatid exchange.

While underlying biological mechanisms responsible for sister chromatid exchange (SCE) formation are not fully understood, scientists worldwide are increasingly using SCEs in the evaluation of excess risk from exposure to chemical and biological agents. SCEs are being used as endpoint measures of cell damage in many types of experimental and nonexperimental investigations. The former includes both simple and complex randomized experiments using both animals exposed in vivo and cells exposed in vitro as experimental units. The latter includes the important, yet potentially misleading, human case-control studies in which a group of humans who are or have been exposed to some agent are compared with a selected nonexposed group on SCE frequency. As more is learned about those factors which result in SCE variations, the assay techniques and study protocols can be adjusted to enhance study sensitivity and to minimize potential bias. Although research concerning sample sizes and statistical analysis methods has been conducted, more is needed. Search for other sources of intersubject variations in SCEs should continue so that such sources can be controlled in future studies, particularly the human exposure types. A number of experimental designs are presented and contrasted with their nonexperimental counterparts. Statistical methods are summarized and sample size options are given for the human and animal exposure studies and for the studies in which cells are exposed in vitro.

Analysis of Variance↗

Statistical phylogeography: methods of evaluating and minimizing inference errors.

Nested clade phylogeographical analysis (NCPA) has become a common tool in intraspecific phylogeography. To evaluate the validity of its inferences, NCPA was applied to actual data sets with 150 strong a priori expectations, the majority of which had not been analysed previously by NCPA. NCPA did well overall, but it sometimes failed to detect an expected event and less commonly resulted in a false positive. An examination of these errors suggested some alterations in the NCPA inference key, and these modifications reduce the incidence of false positives at the cost of a slight reduction in power. Moreover, NCPA does equally well in inferring events regardless of the presence or absence of other, unrelated events. A reanalysis of some recent computer simulations that are seemingly discordant with these results revealed that NCPA performed appropriately in these simulated samples and was not prone to a high rate of false positives under sampling assumptions that typify real data sets. NCPA makes a posteriori use of an explicit inference key for biological interpretation after statistical hypothesis testing. Alternatives to NCPA that claim that biological inference emerges directly from statistical testing are shown in fact to use an a priori inference key, albeit implicitly. It is argued that the a priori and a posteriori approaches to intraspecific phylogeography are complementary, not contradictory. Finally, cross-validation using multiple DNA regions is shown to be a powerful method of minimizing inference errors. A likelihood ratio hypothesis testing framework has been developed that allows testing of phylogeographical hypotheses, extends NCPA to testing specific hypotheses not within the formal inference key (such as the out-of-Africa replacement hypothesis of recent human evolution) and integrates intra- and interspecific phylogeographical inference.

Computer Simulation↗

Detecting activations in PET and fMRI: levels of inference and power.

This paper is about detecting activations in statistical parametric maps and considers the relative sensitivity of a nested hierarchy of tests that we have framed in terms of the level of inference (voxel level, cluster level, and set level). These tests are based on the probability of obtaining c, or more, clusters with k, or more, voxels, above a threshold u. This probability has a reasonably simple form and is derived using distributional approximations from the theory of Gaussian fields. The most important contribution of this work is the notion of set-level inference. Set-level inference refers to the statistical inference that the number of clusters comprising an observed activation profile is highly unlikely to have occurred by chance. This inference pertains to the set of activations reaching criteria and represents a new way of assigning P values to distributed effects. Cluster-level inferences are a special case of set-level inferences, which obtain when the number of clusters c = 1. Similarly voxel-level inferences are special cases of cluster-level inferences that result when the cluster can be very small (i.e., k = 0). Using a theoretical power analysis of distributed activations, we observed that set-level inferences are generally more powerful than cluster-level inferences and that cluster-level inferences are generally more powerful than voxel-level inferences. The price paid for this increased sensitivity is reduced localizing power: Voxel-level tests permit individual voxels to be identified as significant, whereas cluster-and set-level inferences only allow clusters or sets of clusters to be so identified. For all levels of inference the spatial size of the underlying signal f (relative to resolution) determines the most powerful thresholds to adopt. For set-level inferences if f is large (e.g., fMRI) then the optimum extent threshold should be greater than the expected number of voxels for each cluster. If f is small (e.g., PET) the extent threshold should be small. We envisage that set-level inferences will find a role in making statistical inferences about distributed activations, particularly in fMRI.

Brain↗

Extreme between-study homogeneity in meta-analyses could offer useful insights.

OBJECTIVES: Meta-analyses are routinely evaluated for the presence of large between-study heterogeneity. We examined whether it is also important to probe whether there is extreme between-study homogeneity. STUDY DESIGN: We used heterogeneity tests with left-sided statistical significance for inference and developed a Monte Carlo simulation test for testing extreme homogeneity in risk ratios across studies, using the empiric distribution of the summary risk ratio and heterogeneity statistic. A left-sided P=0.01 threshold was set for claiming extreme homogeneity to minimize type I error. RESULTS: Among 11,803 meta-analyses with binary contrasts from the Cochrane Library, 143 (1.21%) had left-sided P-value <0.01 for the asymptotic Q statistic and 1,004 (8.50%) had left-sided P-value <0.10. The frequency of extreme between-study homogeneity did not depend on the number of studies in the meta-analyses. We identified examples where extreme between-study homogeneity (left-sided P-value <0.01) could result from various possibilities beyond chance. These included inappropriate statistical inference (asymptotic vs. Monte Carlo), use of a specific effect metric, correlated data or stratification using strong predictors of outcome, and biases and potential fraud. CONCLUSION: Extreme between-study homogeneity may provide useful insights about a meta-analysis and its constituent studies.

Databases, Bibliographic↗

R.A. Fisher's contributions to genetical statistics.

R. A. Fisher (1890-1962) was a professor of genetics, and many of his statistical innovations found expression in the development of methodology in statistical genetics. However, whereas his contributions in mathematical statistics are easily identified, in population genetics he shares his preeminence with Sewall Wright (1889-1988) and J. B. S. Haldane (1892-1965). This paper traces some of Fisher's major contributions to the foundations of statistical genetics, and his interactions with Wright and with Haldane which contributed to the development of the subject. With modern technology, both statistical methodology and genetic data are changing. Nonetheless much of Fisher's work remains relevant, and may even serve as a foundation for future research in the statistical analysis of DNA data. For Fisher's work reflects his view of the role of statistics in scientific inference, expressed in 1949: There is no wide or urgent demand for people who will define methods of proof in set theory in the name of improving mathematical statistics. There is a widespread and urgent demand for mathematicians who understand that branch of mathematics known as theoretical statistics, but who are capable also of recognising situations in the real world to which such mathematics is applicable. In recognising features of the real world to which his models and analyses should be applicable, Fisher laid a lasting foundation for statistical inference in genetic analyses.

Animals↗