Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Statistical inference by confidence intervals: issues of interpretation and utilization.

This article examines the role of the confidence interval (CI) in statistical inference and its advantages over conventional hypothesis testing, particularly when data are applied in the context of clinical practice. A CI provides a range of population values with which a sample statistic is consistent at a given level of confidence (usually 95%). Conventional hypothesis testing serves to either reject or retain a null hypothesis. A CI, while also functioning as a hypothesis test, provides additional information on the variability of an observed sample statistic (ie, its precision) and on its probable relationship to the value of this statistic in the population from which the sample was drawn (ie, its accuracy). Thus, the CI focuses attention on the magnitude and the probability of a treatment or other effect. It thereby assists in determining the clinical usefulness and importance of, as well as the statistical significance of, findings. The CI is appropriate for both parametric and nonparametric analyses and for both individual studies and aggregated data in meta-analyses. It is recommended that, when inferential statistical analysis is performed, CIs should accompany point estimates and conventional hypothesis tests wherever possible.

Bias↗

Psychological stress and infertility. Part 2: Psychometric test data.

The hypothesis tested in this study was that a group of female functional infertile patients would show significantly more personality maladjustment than a group with definite organic reproductive pathology and a normal fertile group. No significant differences between the functional and organic groups were found on any of the subscores of four personality questionnaires. The normal group (wives of sterile men) scored higher on extraversion than both the functional and organic groups. On self-control the normal group also scored lower (negative connotation) than the other two groups. In general no evidence for personality maladjustment in functional infertility was found.

Adaptation, Psychological↗

Alternative theories of concept identification among older adults.

The purpose of the present experiment was to investigate some predictions of hypothesis testing and S-R association (frequency) theories regarding memory for intratrial events on a conjunctive concept-identification task. They have received extensive study with young adults but not with older subjects. The individual events under investigation were feedback, responses, hypotheses, and stimuli. Hypothesis-testing theory requires subjects to retain information concerning the correct hypothesis from one trial to the next whereas frequency does not. 75 subjects (60-70 yr. old) participated in the study. Subjects had difficulty in recalling the correct hypothesis stated on previous trials. These errors occurred on problems with negative response trials, not with incorrect feedback. The results contradict predictions based on hypothesis-testing models but are consistent with frequency theory. Unlike in the studies based on younger adults, present subjects did not recall the hypothesis very well under the conditions in which hypothesis testing was made part of the primary task.

Aged↗

An efficient procedure for permutation tests in imaging research.

Recent interest in hypothesis testing on functional imaging data has spurred the development of several statistical techniques. The purpose of this paper is to provide a method to reduce the computational intensity associated with randomization tests of positron emission tomography imaging data. We discuss the advantages and disadvantages of traditional distributional hypothesis testing versus the advantages and disadvantages of randomization tests. A method for reducing the computational intensity of randomization uses a conjunction of updating and sequenching and results in significantly reduced processing. The running times of randomization methods are compared.

Algorithms↗

Atenolol in seasonal affective disorder: a test of the melatonin hypothesis.

To test the hypothesis that the antidepressant effects of bright light in seasonal affective disorder are mediated by the suppression of melatonin, 19 patients with this disorder were given atenolol, which suppresses melatonin secretion, and placebo in a double-blind crossover study. No difference in antidepressant efficacy was found between drug and placebo in the sample as a whole, which argues against the melatonin hypothesis of phototherapy. However, in three of the patients atenolol provided repeated, marked, and sustained relief of symptoms, suggesting that it may be useful in treating the winter depressive symptoms of some patients with seasonal affective disorder.

Adult↗

Clinical evaluation of algorithms for ST measurement during exercise test.

HYPOTHESIS: Computer processing of the exercise electrocardiogram (ECG) has many advantages, but the reliability of the analysis algorithms is not easily evaluable. No standard annotated database, nor recommended practice for testing and reporting performance results is available: thus, performance evaluation of such devices can be accomplished only by using a set of unannotated recordings, obtained in clinical practice. We evaluated the accuracy of an original microcomputer-based exercise test analyzer comparing the ST computer output with the measurements obtained by two experienced cardiologists. METHODS: Six hundred ECG strips were randomly selected from the exercise test recordings of 60 patients. The ST shift (at J + 80 ms) was blindly assessed by two observers (with the aid of a calibrated lens) and compared with computer measurements. Correlation coefficients, linear regression equations, percent of discrepant measurements, and 95% confidence limits of the mean error were calculated for all leads, peripheral leads, precordial leads, and "stress-test" leads (II, III, aVF, V4, V5, V6). RESULTS: The computer did not analyze five samples on a total of 600 (0.83%) ECG strips because of excessive noise or signal loss, while 51 (8.5%) were considered unreadable by both observers and 67 (11.2%) were rejected by at least one observer. Correlation between the measurements taken by computer and observer(s) measurements was statistically significant (p < 0.001 for all lead groups), no systematic measurement bias was found, and the mean difference was lower than human eye resolution. CONCLUSIONS: Our algorithms provide results as good as those provided by trained cardiologists in measuring ST changes occurring during exercise test. However, this study did not evaluate whether computer improvement of the signal-to-noise ratio would allow accurate measurements even on cardiologists' uninterpretable ECG. This potential advantage of computer-assisted analysis could be assessed only by using a dedicated exercise test database, in which different patterns of noise are superimposed on noise-free recordings previously annotated for ST level.

Adult↗

General slowing in language impairment: methodological considerations in testing the hypothesis.

Although the general slowing hypothesis of language impairment (LI) is well established, the conventional method to test the hypothesis is controversial. This paper compares the usual method, ordinary least squares regression (OLS), with another method, hierarchical linear modeling with random coefficients (HLM). The analyses used available response time (RT) data from studies of perceptual-motor, cognitive, and language skills of LI and chronological-age-matched (CA) groups. The data set included RT measures from 25 studies investigating 20 different tasks (e.g., auditory detection, mental rotation, and word recognition tasks). OLS and HLM analyses of the RT data yielded very different results. OLS supported general slowing for the LI groups, and indicated that they were significantly slower than CA groups across studies by an overall estimate of 10%. HLM indicated a larger average extent of LI slowing (18%). However, the variability around this average was much greater than that yielded by OLS, and the extent of slowing was not statistically significant. Importantly, HLM showed a significant difference in the RT relation between LI and CA groups across studies, indicating that study-specific slowing, rather than general slowing across studies, was present. A separate HLM analysis of two types of language tasks, picture naming and word recognition, was performed. Although the extent of slowing was equivalent across these tasks, the slowing was minimal (2%) and not significant. Methodological limitations of each analysis to assess general slowing are highlighted.

Adolescent↗

Excess positive associations in communities of intestinal helminths of bats: a refined null hypothesis and a test of the facilitation hypothesis.

The null hypothesis that the number of positive pairwise covariances should equal the number of negative pairwise covariances in samples from communities of randomly associated helminth species was reevaluated. The proportion of positive covariances in a sample from a community of independent species depends upon the proportion of rare species (prevalence less than 10%), the proportion of common species (prevalence greater than 90%), and the size of the sample of hosts. If rare species dominate, then there will be an excess of negative associations; if common species dominate there will be an excess of positive associations. Many helminth communities have more rare than common species, therefore samples from communities that show an equal number of positive and negative covariances have a greater number of positive associations than is expected for randomly associated species. Increased sample size will reduce the sampling bias, but at least 100 hosts are necessary and often 500-7,500 hosts are required. The excess of positive covariances between helminth species in 10 populations of bats disappeared after restricting the analyses to hosts in which both members of a species pair were present. This result suggests that excess positive associations between helminth species in bats are due to joint presences and absences in hosts rather than to interspecific facilitation. Interspecific facilitation would be supported by observed positive correlations between the intensities of individuals of the species pairs.

Analysis of Variance↗

[Principles of tests of hypotheses in statistics: alpha, beta and P].

Modern clinical research requires control of statistical methods. We reviewed 120 original manuscripts which were submitted to the Annales françaises d'anesthésie et de reanimation and analyzed their statistical methodology. Most of them contained errors (inappropriate numerical expression of the data, uncontrolled alpha risk, lack of power, use of inadequate statistical tests) and only 9 (7%) were considered as adequate. Therefore it is useful to come back to the methodology of hypothesis testing. An hypothesis test helps to decide between two hypothesis, the null hypothesis (H0) and the alternative hypotheses (H1) that we intend to demonstrate. The decision of the choice between H0 and H1 is associated with two probabilities: the alpha risk which is the probability to reject H0 whereas H0 is true, and the beta risk which is the probability not to reject H0 whereas H1 is true. Because the alpha risk is considered to be very important, it should be verified that the actual risk corresponds to the risk initially retained. The P value is the probability to observe a difference as great as that noted. The P value should be assessed according to its environment: the clinical relevance of a result should be assessed according to the amplitude of the difference and its confidence interval. When the null hypothesis is not rejected, the power of the test is essential. Power calculation is essential in clinical research trials. The number of patients included depends on four elements: the response to the control treatment, the expected response to the new treatment, the level of significance, and the power. The following items should be checked to choose the appropriate test: assess the kind of variable, verify the requirements for application of the test (type of the variable distribution, sample size, particular conditions such as equality of variance, dependence or independence of the variables), determine if data come from paired samples or if multiple comparisons are performed. Statistical analysis has become more easy with computers, however a precise knowledge of statistics remains essential. Advice from a statistician is often useful especially when obtained a priori and not a posteriori.

Analysis of Variance↗

A unified method for monitoring and analysing controlled trials.

Group sequential methods are becoming increasingly popular for monitoring and analysing large controlled trials, especially clinical trials. They not only allow trialists to monitor the data as it accumulates, but also reduce the expected sample size. Such methods are traditionally based on preserving the overall type I error by increasing the conservatism of the hypothesis tests performed at any single analysis. Using methods which are based on hypothesis testing in this way makes point estimation and the calculation of confidence intervals difficult and controversial. We describe a class of group sequential procedures based on a single parameter which reflects initial scepticism towards unexpectedly large effects. These procedures have good expected and maximum sample sizes, and lead to natural point and interval estimates of the treatment difference. Hypothesis tests, point estimates and interval estimates calculated using this procedure are consistent with each other, and tests and estimates made at the end of the trial are consistent with interim tests and estimates. This class of sequential tests can be considered in both a traditional group sequential manner or as a Bayesian solution to the problem.

Bayes Theorem↗

Reasoning biases in delusion-prone individuals.

OBJECTIVES: The objective was to test whether individuals high in delusional ideation exhibit a reasoning bias on tasks involving hypothesis testing and probability judgments. On the basis of previous findings (e.g. Garety, Hemsley & Wessely, 1991), it was predicted that individuals high in delusional ideation would exhibit a 'jump-to-conclusions' style of reasoning and would be less sensitive to the effects of random variation, in comparison to individuals low in delusional ideation. DESIGN: A non-randomized matched groups design was employed enabling the performance of the delusion prone individuals to be compared to that of a control group. METHOD: Forty individuals, selected from the normal population, were divided into groups high and low in delusional ideation, according to their scores on the Peters et al. Delusions Inventory (Peters, Day & Garety, 1996), and were compared on two tasks involving probability judgment and two tasks involving hypothesis testing. RESULTS: Although no significant differences were found on tasks involving hypothesis testing and the aggregation of probabilistic information, it was found that individuals high in delusional ideation had a 'jump-to-conclusions' style of data gathering and were less sensitive to the effects of random variation, in comparison to individuals low in delusional ideation. CONCLUSIONS: In conclusion, although individuals high in delusional ideation were not found to have a general reasoning bias, some evidence of a more specific bias was found. It is thought that these aberrations may play some role in delusion formation in schizophrenia and paranoia.

Adult↗

Testing the hypothesis of common ancestry.

The hypothesis that all life on earth traces back to a single common ancestor is a fundamental postulate in modern evolutionary theory. Yet, despite its widespread acceptance in biology, there has been comparatively little attention to formally testing this "hypothesis of common ancestry". We review and critically examine some arguments that have been proposed in support of this hypothesis. We then describe some theoretical results that suggest the hypothesis may be intrinsically difficult to test. We conclude by suggesting an approach to the problem based on the Aikaike information criterion.

Animals↗

Importance of using rigorous statistical methods to analyze low energy laser experimental data: Part two.

BACKGROUND AND OBJECTIVE: Numerous authors have reported successful alteration of peripheral nerve action potential characteristics through application of low energy laser irradiation (LELI). The statistical analysis that accompanies many of these reports frequently does not account for the special nature of the data generated in typical LELI experiments. The objective of this study was to evaluate the application of repeated measures linear regression techniques to the analysis of this type of data. Issues of analyzing raw versus normalized data, proper accounting for correlation between measurements, and discrete time point hypothesis testing were addressed. STUDY DESIGN/MATERIALS AND METHODS: The data analyzed in this work were generated from an experiment in which in vitro frog sciatic nerves were irradiated with a helium-neon laser using a variety of treatment protocols. Compound action potential (CAP) amplitude, latency, depolarization rate, and repolarization rate were recorded at 1-minute intervals for 135 minutes for each nerve. Laser-induced changes in CAP parameters were analyzed using various repeated measures linear regression models. RESULTS: The findings of statistical significance were highly dependent on the rigor of the regression model applied. Application of the same regression model to raw and normalized data produced different findings of significance. Determination of significant contrasts was highly dependent on how well the regression model accounted for the correlation between repeated measurements made on the same nerve. In general, models that failed to account adequately for this correlation produced more findings of significant contrasts than increasingly rigorous models. Finally, discrete time point hypothesis testing on normalized data can suggest improper statistical conclusions if the proper correlation structure is not applied to the data set. CONCLUSION: Linear regression analysis offers advantages over discrete time point hypothesis testing in the analysis of highly correlated serial data of this type. Trends in the behavior of the measured parameters are evident, rigorous accounting for correlation between measurements is facilitated, and hypothesis testing is highly flexible.

Action Potentials↗

An automated method for neuroanatomic and cytoarchitectonic atlas-based interrogation of fMRI data sets.

Analysis and interpretation of functional MRI (fMRI) data have traditionally been based on identifying areas of significance on a thresholded statistical map of the entire imaged brain volume. This form of analysis can be likened to a "fishing expedition." As we become more knowledgeable about the structure-function relationships of different brain regions, tools for a priori hypothesis testing are needed. These tools must be able to generate region of interest masks for a priori hypothesis testing consistently and with minimal effort. Current tools that generate region of interest masks required for a priori hypothesis testing can be time-consuming and are often laboratory specific. In this paper we demonstrate a method of hypothesis-driven data analysis using an automated atlas-based masking technique. We provide a powerful method of probing fMRI data using automatically generated masks based on lobar anatomy, cortical and subcortical anatomy, and Brodmann areas. Hemisphere, lobar, anatomic label, tissue type, and Brodmann area atlases were generated in MNI space based on the Talairach Daemon. Additionally, we interfaced these multivolume atlases to a widely used fMRI software package, SPM99, and demonstrate the use of the atlas tool with representative fMRI data. This tool represents a necessary evolution in fMRI data analysis for testing of more spatially complex hypotheses.

Brain↗

Reliability of measurement of angular movements of the pelvis and lumbar spine during treadmill walking.

BACKGROUND AND PURPOSE: Angular movements of the pelvis and lumbar spine are thought to play an important role in walking. However, little is known about the amount of unpredictable variability in measurement of these movements during human walking. The aim of the present study was to determine the retest reliability of measuring the angular movements of the pelvis and lumbar spine during unimpaired familiarized treadmill walking. METHOD: Retest reliability for 26 subjects without pathology was determined over a one-week interval. Subjects walked on a treadmill at self-selected or a slower speed while measurements of the three-dimensional angular movements were taken with a computer-based video analysis system. RESULTS: The frontal plane movements of pelvic list and lumbar lateral flexion (relative to the pelvis) could be measured with high retest reliability at both self-selected and slow walking speeds (intraclass coefficient (ICC) (2, 1) > or = 0.81). In contrast, transverse and sagittal plane movements demonstrated moderate reliability at both speeds (0.37 < or = ICC (2, 1) < or = 0.76). Averaging the measurement over six strides resulted in increased observed reliability (self-selected walking speed summary Pearson's r = 0.71, slow walking speed summary Pearson's r = 0.79) compared to taking the measurement based on a single stride (self-selected walking speed summary Pearson's r = 0.63, slow walking speed summary Pearson's r = 0.67). Unlike pelvic and lumbar movements (relative to the pelvis), the measurement of lumbar movements (relative to the global reference frame) appeared to depend on whether subjects were walking at self-selected or slow speeds. CONCLUSIONS: Measurement of pelvic list and lumbar lateral flexion (relative to the pelvis) could be applied with confidence to hypothesis testing about individuals or groups. Movements in the transverse and sagittal planes are unlikely to be appropriate in hypothesis testing about individuals and hence clinical practice, but may still have experimental applications in hypothesis testing about groups.

Adult↗

Group comparisons involving missing data in clinical trials: a comparison of estimates and power (size) for some simple approaches.

When using 'intent-to-treat' approaches to compare outcomes between groups in clinical trials, analysts face a decision regarding how to account for missing observations. Most model-based approaches can be summarized as a process whereby the analyst makes assumptions about the distribution of the missing data in an attempt to obtain unbiased estimates that are based on functions of the observed data. Although pointed out by Rubin as often leading to biased estimates of variances, an alternative approach that continues to appear in the applied literature is to use fixed-value imputation of means for missing observations. The purpose of this paper is to provide illustrations of how several fixed-value mean imputation schemes can be formulated in terms of general linear models that characterize the means of distributions of missing observations in terms of the means of the distributions of observed data. We show that several fixed-value imputation strategies will result in estimated intervention effects that correspond to maximum likelihood estimates obtained under analogous assumptions. If the missing data process has been correctly characterized, hypothesis tests based on variances estimated using maximum likelihood techniques asymptotically have the correct size. In contrast, hypothesis tests performed using the uncorrected variance, obtained by applying standard complete data formula to singly imputed data, can provide either conservative or anticonservative results. Surprisingly, under several non-ignorable non-response scenarios, maximum likelihood based analyses can yield equivalent hypothesis tests to those obtained when analysing only the observed data.

Aged↗

Maximizing outcomes while minimizing exploration in hyperparathyroidism using localization tests.

HYPOTHESIS: Preoperative localization (ultrasonography and scintigraphy) can be used to limit operative exploration in primary hyperparathyroidism while providing a high rate of success. DESIGN: Prospective cohort analysis of 3 types of exploration (1-gland, unilateral, or 4-gland), as directed by localization. RESULTS: In 185 consecutive patients who underwent operations, the final diagnoses were solitary adenoma in 87% and multigland disease in 13%. Ultrasonography (75%) and scintigraphy (83%) demonstrated an enlarged parathyroid gland and, together with operative findings, resulted in 61 1-gland, 63 unilateral, and 61 4-gland explorations, with an initial success rate of 96% and an ultimate success rate of 99%. Limiting exploration resulted in a significant decrease in operative time and hospitalization. CONCLUSION: Localization can limit exploration with success.

Adenoma↗

Testing the hypothesis that system y(+)L accounts for high- and low-transport phenotypes in chicken erythrocytes using L-leucine as substrate.

Experiments were carried out to test the hypothesis that system y(+)L accounts for the high (HT) and low (LT) amino-acid transport phenotypes in chicken erythrocytes and to explain the different effect of selective breeding on lysine and leucine fluxes. L: -Leucine transport was characterized in individuals which had been separated into two groups (HT and LT) according to their capacity to transport L: -lysine across the erythrocyte membrane. Whereas lysine influx (1 muM: ) in the two groups differed by 32-fold (HT/LT), leucine influx was not significantly different. Average rates (nmol/ L cells/ min) were: 227 (HT) and 7.0 (LT) for L: -lysine, and 8.9 (HT) and 7.1 (LT) for L: -leucine. The differential ability of L: -lysine and L: -leucine fluxes to discriminate between the HT and LT phenotypes was shown to be consistent with the interactions of these substrates with system y(+)L and to vary depending on the conditions of the assay. It is shown that the two phenotypes can be clearly discriminated by measuring L: -leucine influx in the presence of Li(+). These results support the hypothesis that the HT and LT phenotypes reflect alterations in the function of system y(+)L and illustrate that the choice of the appropriate substrate and medium composition must be carefully considered when investigating the consequences of either experimental or natural alterations of broad-scope transporters.

Amino Acid Transport System y+L↗