Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Statistical reasoning in clinical trials: hypothesis testing.

Hypothesis testing is based on certain statistical and mathematical principles that allow investigators to evaluate data by making decisions based on the probability or implausibility of observing the results obtained. However, classic hypothesis testing has its limitations, and probabilities mathematically calculated are inextricably linked to sample size. Furthermore, the meaning of the p value frequently is misconstrued as indicating that the findings are also of clinical significance. Finally, hypothesis testing allows for four possible outcomes, two of which are errors that can lead to erroneous adoption of certain hypotheses: 1. The null hypothesis is rejected when, in fact, it is false. 2. The null hypothesis is rejected when, in fact, it is true (type I or alpha error). 3. The null hypothesis is conceded when, in fact, it is true. 4. The null hypothesis is conceded when, in fact, it is false (type II or beta error). The implications of these errors, their relation to sample size, the interpretation of negative trials, and strategies related to the planning of clinical trials will be explored in a future article in this journal.

Clinical Trials as Topic↗

Hypothesis testing.

Hypothesis testing is the process of making a choice between two conflicting hypotheses. The null hypothesis, H0, is a statistical proposition stating that there is no significant difference between a hypothesized value of a population parameter and its value estimated from a sample drawn from that population. The alternative hypothesis, H1 or Ha, is a statistical proposition stating that there is a significant difference between a hypothesized value of a population parameter and its estimated value. When the null hypothesis is tested, a decision is either correct or incorrect. An incorrect decision can be made in two ways: We can reject the null hypothesis when it is true (Type I error) or we can fail to reject the null hypothesis when it is false (Type II error). The probability of making Type I and Type II errors is designated by alpha and beta, respectively. The smallest observed significance level for which the null hypothesis would be rejected is referred to as the p-value. The p-value only has meaning as a measure of confidence when the decision is to reject the null hypothesis. It has no meaning when the decision is that the null hypothesis is true.

Bias↗

A shift from significance test to hypothesis test through power analysis in medical research.

Medical research literature until recently, exhibited substantial dominance of the Fisher's significance test approach of statistical inference concentrating more on probability of type I error over Neyman-Pearson's hypothesis test considering both probability of type I and II error. Fisher's approach dichotomises results into significant or not significant results with a P value. The Neyman-Pearson's approach talks of acceptance or rejection of null hypothesis. Based on the same theory these two approaches deal with same objective and conclude in their own way. The advancement in computing techniques and availability of statistical software have resulted in increasing application of power calculations in medical research and thereby reporting the result of significance tests in the light of power of the test also. Significance test approach, when it incorporates power analysis contains the essence of hypothesis test approach. It may be safely argued that rising application of power analysis in medical research may have initiated a shift from Fisher's significance test to Neyman-Pearson's hypothesis test procedure.

Biomedical Research↗

Statistical hypothesis testing--how exact are exact p-values?

OBJECTIVES AND BACKGROUND: When testing a hypothesis statistically, a principle is generally accepted that exact p values shall be stated in the treatise. Researchers have the choice of many statistical computer programmes with implemented hypothesis tests. Are exact p values calculated in the same statistical tests by diverse statistical programmes identical? METHODS: The respective zero hypothesis were tested in 5 artificially created data sets by the parametric unpaired t-test, non-parametric Mann-Whitney test, two-tailed F-test. The calculations were carried out by the following programmes: Statistix, version 7.1 (source www.statistix.com), Analyse-it, version 1.62 (source www.analyse-it.com), MedCalc, version 6.14 (source www.medcalc.be). The p values in the same tests were mutually compared. RESULTS: All three programmes calculated identical exact p values for the t-test. In the remaining two tests in case of 26 out of 44 calculations (59.1 per cent; 95 per cent confidence interval 43-73 per cent) different p values were calculated. The greatest difference was 18.35 per cent. In two cases the values oscillated about 0.05 and this fact caused essentially different interpretation of results. CONCLUSIONS: Using the significance test in the biomedical research has been subject to criticism for a longer period of time. The testing of the zero hypothesis on the arbitrary significance level of 0.05 should be substituted by other methods. Our discoveries should undermine the ungrounded belief of the users of statistical tests--physicians in ununderminable accuracy of mathematical procedures. The use of confidence intervals deems much more suitable although there are objections against them as well. (Tab. 4, Fig. 1, Ref. 19.).

Confidence Intervals↗

Hypothesis testing and anomaly explaining.

An experimenter who tests a hypothesis and observes an anomaly that conflicts with his knowledge and views, normally reacts by suspecting some sort of mistake. When he has excluded this possibility, his second sensible reaction is to find an ad hoc interpretation to explain the anomaly. When his favorite interpretation is interesting, probable, simple, elegant and testable enough, he can experiment to verify this explanation as a new hypothesis, independently from the anomaly. This article describes how a scientific investigation walks on two legs: one leg of hypothesis testing and a second leg of anomaly explaining.

Bias↗

Confidence intervals rather than P values: estimation rather than hypothesis testing.

Overemphasis on hypothesis testing--and the use of P values to dichotomise significant or non-significant results--has detracted from more useful approaches to interpreting study results, such as estimation and confidence intervals. In medical studies investigators are usually interested in determining the size of difference of a measured outcome between groups, rather than a simple indication of whether or not it is statistically significant. Confidence intervals present a range of values, on the basis of the sample data, in which the population value for such a difference may lie. Some methods of calculating confidence intervals for means and differences between means are given, with similar information for proportions. The paper also gives suggestions for graphical display. Confidence intervals, if appropriate to the type of study, should be used for major findings in both the main text of a paper and its abstract.

Adult↗

Effectiveness of Positive Hypothesis Testing in Inductive and Deductive Rule Learning.

In a positive hypothesis test a person generates or examines evidence that is expected to have the property of interest if the hypothesis is correct, whereas in a negative hypothesis test a person generates or examines evidence that is not expected to have the property of interest if the hypothesis is correct. Two experiments assessed the effectiveness of positive versus negative hypothesis tests on inductive and deductive rule learning problems. In Experiment 1 problem solvers induced a rule by proposing hypotheses and selecting evidence in the eight conditions of a factorial design defined by instructions to use a positive or negative hypothesis test on each of trials 1-5, 6-10, and 11-15. Instructions to use positive tests resulted in more examples, fewer strategic hypotheses, and a higher weighted score for five types of hypotheses than instructions to use negative tests. In Experiment 2 problem solvers identified 1 of a possible 1296 correct rules in the deductive rule learning game Mastermind. When problems were classified in the 16 possible combinations of positive or negative hypothesis tests on trials 2, 3, 4, and 5 there were fewer trials to solution for positive tests on each of the four trials and fewer trials to solution with increasing positive tests. We conclude that positive hypothesis tests are generally more effective than negative hypothesis tests in both inductive and deductive rule learning. Copyright 1999 Academic Press.

Journal Article↗

[Hypothesis testing processes in the reception strategy task and hypothesis evaluation task].

This study investigated the common characteristics of reasoning in the two types of hypothesis testing tasks that contain similar passive information gathering procedures: reception strategy task and hypothesis evaluation task. Twenty subjects built their own hypothesis but were not allowed to choose instances to test the hypothesis in the reception strategy task, and different 20 subjects in the hypothesis evaluation task only evaluated a hypothesis generated by another person. In addition, subjects were asked to mark instances that they thought suitable to test the hypothesis on each trial. Results showed that subjects chose negative (-H) tests in the middle or later stage of the task. Otherwise, they chose confirmative tests throughout the task. Two possible interpretations of the results were offered that (a) experiencing the false negative (-H +T) made subjects to realize the usefulness of negative testing, and (b) there was a phased shift of the selection tendency between the earlier phase of extracting a possible hypothesis to the later phase of refining it.

Problem Solving↗

Testing for bimodality in frequency distributions of data suggesting polymorphisms of drug metabolism--hypothesis testing.

1. The theory of methods of hypothesis testing in relation to the detection of bimodality in density distributions is discussed. 2. Practical problems arising from these methods are outlined. 3. The power of three methods of hypothesis testing was compared using simulated data from bimodal distributions with varying separation between components. None of the methods could determine bimodality until the separation between components was 2 standard deviation units and could only do so reliably (greater than 90%) when the separation was as great as 4-6 standard deviation units. 4. The robustness of a parametric and a non-parametric method of hypothesis testing was compared using simulated unimodal distributions known to deviate markedly from normality. Both methods had a high frequency of falsely indicating bimodality with distributions where the components had markedly differing variances. 5. A further test of robustness using power transformation of data from a normal distribution showed that the algorithms could accurately determine unimodality only when the skew of the distribution was in the range 0-1.45.

Computers↗

Hypothesis testing in distributed source models for EEG and MEG data.

Hypothesis testing in distributed source models for the electro- or magnetoencephalogram is generally performed for each voxel separately. Derived from the analysis of functional magnetic resonance imaging data, such a statistical parametric map (SPM) ignores the spatial smoothing in hypothesis testing with distributed source models. For example, when intending to test a single voxel, actually an entire region of voxels is tested simultaneously. Because there are more parameters than observations, typically constraints are employed to arrive at a solution which spatially smooths the solution. If ignored, it can be concluded from the hypothesis test that there is activity at some location where there is none. In addition, an SPM on distributed source models gives the illusion of very high resolution. As an alternative, a multivariate approach is suggested in which a region of interest is tested that is spatially smooth. In simulations with MEG and EEG it is shown that clear hypothesis testing in distributed source models is possible, provided that there is high correspondence between what is intended to be tested and what is actually tested. The approach is also illustrated by an application to data from an experiment measuring visual evoked fields when presenting checkerboard patterns.

Brain↗

Hypothesis and hypothesis testing in the clinical trial.

The hypothesis provides the justification for the clinical trial. It is antecedent to the trial and establishes the trial's direction. Hypothesis testing is the most widely employed method of determining whether the outcome of clinical trials is positive or negative. Too often, however, neither the hypothesis nor the statistical information necessary to evaluate outcomes, such as p values and alpha levels, is stated explicitly in reports of clinical trials. This article examines 5 recent studies comparing atypical antipsychotics with special attention to how they approach the hypothesis and hypothesis testing. Alternative approaches are also discussed.

Antipsychotic Agents↗

Hypothesis testing for selective, differential, and conjoined brain activation.

Hypothesis testing in functional neuroimaging studies relies heavily on the computation of categorical contrasts in which brain activation associated with one experimental condition is assessed relative to brain activation associated with a different experimental condition. Often, multiple pair-wise contrasts are computed and reported independently. Here we describe an approach to hypothesis testing that logically combines multiple pair-wise contrasts to distinguish among selective, differential and conjoined brain activation patterns. Using a sample dataset in which participants viewed objects, visual noise patterns or a fixation cross, we demonstrate that selective and differential brain activation patterns are often confounded with current approaches to hypothesis testing but that the logical combination approach can distinguish between these two types of data patterns. Specifically, we show that brain regions that respond selectively to an object recognition task relative to viewing visual noise or a fixation cross (selective activation) are mutually exclusive from brain regions that show a graded response to object viewing, noise viewing and visual fixation (differential activation). We thus show that the logical combination approach sufficiently constrains the results of categorical contrasts to reflect only the data pattern that would be predicted from the cognitive processing account under investigation.

Adolescent↗

On the logic of hypothesis testing in functional imaging.

Statistics is nowadays the customary language of functional imaging. It is common to express an experimental setting as a set of null hypotheses over complex models and to present results as maps of p-values derived from sophisticated probability distributions. However, the growing interest in the development of advanced statistical algorithms is not always paralleled by similar attention to how these techniques may regiment the ways in which users draw inferences from their data. This article investigates the logical bases of current statistical approaches in functional imaging and probes their suitability to inductive inference in neuroscience. The frequentist approach to statistical inference is reviewed with attention to its two main constituents: Fisherian "significance testing" and Neyman-Pearson "hypothesis testing". It is shown that these conceptual systems, which are similar in the univariate testing case, dissociate into two quite different methods of inference when applied to the multiple testing problem, the typical framework of functional imaging. This difference is explained with reference to specific issues, like small volume correction, which are most likely to generate confusion in the practitioner. Further insight into this problem is achieved by recasting the multiple comparison problem into a multivariate Bayesian formulation. This formulation introduces a new perspective where the inferential process is more clearly defined in two distinct steps. The first one, inductive in form, uses exploratory techniques to acquire preliminary notions on the spatial patterns and the signal and noise characteristics. The (smaller) set of likely spatial patterns generated is then tested with newer data and a more rigorous multiple hypothesis testing technique (deductive step).

Algorithms↗

A comparative evaluation of wavelet-based methods for hypothesis testing of brain activation maps.

Wavelet-based methods for hypothesis testing are described and their potential for activation mapping of human functional magnetic resonance imaging (fMRI) data is investigated. In this approach, we emphasise convergence between methods of wavelet thresholding or shrinkage and the problem of hypothesis testing in both classical and Bayesian contexts. Specifically, our interest will be focused on the trade-off between type I probability error control and power dissipation, estimated by the area under the ROC curve. We describe a technique for controlling the false discovery rate at an arbitrary level of error in testing multiple wavelet coefficients generated by a 2D discrete wavelet transform (DWT) of spatial maps of fMRI time series statistics. We also describe and apply change-point detection with recursive hypothesis testing methods that can be used to define a threshold unique to each level and orientation of the 2D-DWT, and Bayesian methods, incorporating a formal model for the anticipated sparseness of wavelet coefficients representing the signal or true image. The sensitivity and type I error control of these algorithms are comparatively evaluated by analysis of "null" images (acquired with the subject at rest) and an experimental data set acquired from five normal volunteers during an event-related finger movement task. We show that all three wavelet-based algorithms have good type I error control (the FDR method being most conservative) and generate plausible brain activation maps (the Bayesian method being most powerful). We also generalise the formal connection between wavelet-based methods for simultaneous multiresolution denoising/hypothesis testing and methods based on monoresolution Gaussian smoothing followed by statistical testing of brain activation maps.

Algorithms↗

Introduction to biostatistics: Part 3, Sensitivity, specificity, predictive value, and hypothesis testing.

Diagnostic tests guide physicians in assessment of clinical disease states, just as statistical tests guide scientists in the testing of scientific hypotheses. Sensitivity and specificity are properties of diagnostic tests and are not predictive of disease in individual patients. Positive and negative predictive values are predictive of disease in patients and are dependent on both the diagnostic test used and the prevalence of disease in the population studied. These concepts are best illustrated by study of a two by two table of possible outcomes of testing, which shows that diagnostic tests may lead to correct or erroneous clinical conclusions. In a similar manner, hypothesis testing may or may not yield correct conclusions. A two by two table of possible outcomes shows that two types of errors in hypothesis testing are possible. One can falsely conclude that a significant difference exists between groups (type I error). The probability of a type I error is alpha. One can falsely conclude that no difference exists between groups (type II error). The probability of a type II error is beta. The consequence and probability of these errors depend on the nature of the research study. Statistical power indicates the ability of a research study to detect a significant difference between populations, when a significant difference truly exists. Power equals 1-beta. Because hypothesis testing yields "yes" or "no" answers, confidence intervals can be calculated to complement the results of hypothesis testing. Finally, just as some abnormal laboratory values can be ignored clinically, some statistical differences may not be relevant clinically.

Biometry↗

Hypothesis testing in ecology: psychological aspects and the importance of theory maturation.

Proper hypothesis testing is the subject of much debate in ecology. According to studies in cognitive psychology, confirmation bias (a tendency to seek confirming evidence) and theory tenacity (persistent belief in a theory in spite of contrary evidence) pervasively influence actual problem solving and hypothesis testing, often interfering with effective testing of alternative hypotheses. On the other hand, these psychological factors play a positive role in the process of theory maturation by helping to protect and nurture a new idea until it is suitable for critical evaluation. As a theory matures it increases in empirical content and its predictions become more distinct. Efficient hypothesis testing is often not possible when theories are in an immature state, as is the case in much of ecology. Problem areas in ecology are examined in light of these considerations, including failure to publish negative results, misuses of mathematical models, confusion resulting from ambiguous terms (such as "diversity" and "niche"), and biases against new ideas.

Cognition↗

Effectiveness of Positive Hypothesis Testing for Cooperative Groups.

In a rule induction problem positive hypothesis tests select evidence that the tester expects to be an example of the correct rule if the hypothesis is correct, whereas negative hypothesis tests select evidence that the tester expects to be a nonexample if the hypothesis is correct. We extend previous analyses of the effectiveness of positive and negative tests for ambiguous verification or conclusive falsification of hypotheses by emphasizing the importance of examples following positive or negative tests. Cooperative four-person groups solved rule induction problems from a single known example of the correct rule by proposing hypotheses and selecting evidence on each of four arrays on a series of trials. There were more examples following positive tests than negative tests. The transition probability from an incorrect hypothesis on trial t to the correct hypothesis on trial t + 1 was higher for positive tests than for negative tests, higher for positive tests followed by examples than positive tests followed by nonexamples, and higher for negative tests followed by examples than negative tests followed by nonexamples. Once the group proposed the correct hypothesis on trial t they were highly likely to continue to propose the correct hypothesis on trial t + 1. Copyright 1998 Academic Press.

Journal Article↗

An empirical analysis of eating disorders and anxiety disorders publications (1980-2000)--part II: Statistical hypothesis testing.

OBJECTIVE: The current study compared the eating disorder literature and the anxiety disorder literature in terms of statistical hypothesis testing features in 1980, 1990, and 2000. METHOD: Computer literature searches were conducted using PubMed and PsychInfo databases to identify relevant eating disorder and anxiety disorder articles published at each of the three time points. A total of 456 articles were randomly selected, including 228 articles each from the fields of eating disorders and anxiety disorders. Within each field, one third (76) of the articles were selected from each of the three time points. Two raters, from a team of eight trained raters, were randomly assigned to independently rate each article in terms of 75 separate methodologic features. In the current article, we will emphasize the findings about hypothesis testing and statistical analysis. Disagreements in ratings were resolved via consensus. Ratings were tabulated separately by field across the three time points. RESULTS: Few differences were observed between eating disorder and anxiety disorder publications in terms of statistical hypothesis testing features. Although increases were observed in both fields in a number of areas from 1980 to 2000, there remains a pervasive absence of many of the statistical hypothesis testing features recommended by the American Psychological Association Task Force on Statistical Inference. CONCLUSION: These results are discussed in terms of their implications for the fields of eating disorders and anxiety disorders, for researchers, for reviewers, and for professional journals and editorial boards.

Anxiety Disorders↗