Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

[What does "p" mean at conclusion of a test of hypothesis in a randomized controlled clinical trial of superiority?].

The aim of this statistical note, the sixth in the series, is to introduce the rationale of the test of hypothesis suitable for comparing the effect of two treatments in a randomized controlled clinical trial of superiority. The presentation takes advantage of the analogy with a criminal trial debate based upon circumstantial evidence in an Italian Court. The results of three randomized controlled clinical trials: ISIS-1, AIMS and RESTORE are introduced and proper ways for their interpretation are suggested.

Randomized Controlled Trials as Topic↗

Simultaneous inference for generalized linear models with unmeasured confounders.

Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may be substantially biased. This paper investigates the large-scale hypothesis testing problem for multivariate generalized linear models in the presence of confounding effects. Under arbitrary confounding mechanisms, we propose a unified statistical estimation and inference framework that harnesses orthogonal structures and integrates linear projections into three key stages. It begins by disentangling marginal and uncorrelated confounding effects to recover the latent coefficients. Subsequently, latent factors and primary effects are jointly estimated through lasso-type optimization. Finally, we incorporate projected and weighted bias-correction steps for hypothesis testing. Theoretically, we establish the identification conditions of various effects and non-asymptotic error bounds. We show effective Type-I error control of asymptotic-tests as sample and response sizes approach infinity. Numerical experiments demonstrate that the proposed method controls the false discovery rate by the Benjamini-Hochberg procedure and is more powerful than alternative methods. By comparing single-cell RNA-seq counts from two groups of samples, we demonstrate the suitability of adjusting confounding effects when significant covariates are absent from the model.

Hidden variables↗

Alcohol-abusing teenage boys. Testing a hypothesis on the relationship between alcohol abuse and social background factors, criminality and personality in teenage boys.

A study material of teenage boys from the general population was used to test the hypothesis on early alcohol abuse suggested by the results of previous prospective studies on selected materials. The results of interviews of 1,004 18-year-old boys from the general population in the Stockholm area support the hypothesis that there is a group (about 4% during autumn 1980) of boys with a high consumption of alcohol, simultaneous use of drugs and criminal behaviour. As a group these boys had been brought up in emotionally disturbed homes, with alcoholic parents, and they also showed personality features indicating psychopathy. The study provides evidence that results of investigations on selected materials are also relevant to the general population.

Adolescent↗

Pros and cons of permutation tests in clinical trials.

Hypothesis testing, in which the null hypothesis specifies no difference between treatment groups, is an important tool in the assessment of new medical interventions. For randomized clinical trials, permutation tests that reflect the actual randomization are design-based analyses for such hypotheses. This means that only such design-based permutation tests can ensure internal validity, without which external validity is irrelevant. However, because of the conservatism of permutation tests, the virtues of permutation tests continue to be debated in the literature, and conclusions are generally of the type that permutation tests should always be used or permutation tests should never be used. A better conclusion might be that there are situations in which permutation tests should be used, and other situations in which permutation tests should not be used. This approach opens the door to broader agreement, but begs the obvious question of when to use permutation tests. We consider this issue from a variety of perspectives, and conclude that permutation tests are ideal to study efficacy in a randomized clinical trial which compares, in a heterogeneous patient population, two or more treatments, each of which may be most effective in some patients, when the primary analysis does not adjust for covariates. We propose the p-value interval as a novel measure of the conservatism of a permutation test that can be defined independently of the significance level. This p-value interval can be used to ensure that the permutation test have both good global power and an acceptable degree of conservatism.

Humans↗

An examination of graduate students' statistical judgments: statistical and fuzzy set approaches.

The present study examined how statistical significance levels are treated and interpreted by graduate students who use hypothesis-testing in their scientific investigation. To test underlying psychological aspects of hypothesis-testing, the idea of fuzzy set theory was employed to identify the uncertain points in judgments. 34 graduate students in a psychology department made judgments about hypothetical statistical decisions. The results indicated that (1) the majority of these students treated significance levels on a continuum and rated them according to the magnitude of statistical significance; (2) the subjects shifted their decisions based on the types of hypothetical scenarios but not by the sample sizes; instead, they interpreted a smaller sample size as being less reliable. (3) The subjects frequently chose formally used statistical terms, e.g., Significant and Not Significant, more than graduated verbal expressions, e.g., Marginally Significant and Borderline Significant; and (4) the Fuzziness (degree of confidence in decision-making) was dependent on individuals and existed more in the critical points of transition where judgments are most difficult. The Fuzziness Index illustrated the subtle shifts of human decision-making patterns in statistical judgments. Underlying decision uncertainties and difficulties can be illustrated by functions generated from fuzzy set theory, which may more closely resemble human psychological mechanism. This integrative study of fuzzy set theory and behavioral measurements appears to provide a technique that is more natural for examining and understanding imprecise boundaries of human decisions.

Adult↗

Testing the hypothesis that crypt size determines the rate of enterocyte development in neonatal mice.

Pieces of mid-jejunum taken from 7-9 day old mice have been used to determine microvillus length in enterocytes located at different points along the crypt-villus axis to test the hypothesis that enterocyte development of structure is directly determined by the physical characteristics of the intestinal crypt. Parallel measurements of enterocyte migration rate were carried out using tritiated thymidine to determine the time course of microvillus elongation in neonatal mice. Microvillus length approximately doubled during early enterocyte migration from the crypt base to the lower part of the villus. Enterocyte migration rate was only 0.9 micron/hr at this stage of development, a value considerably less than that found in adult intestine. Results plotting the time dependency of microvillus elongation were fitted by a logistic curve giving a maximal rate for microvillus growth of 0.004 micron/hr. The corresponding estimate of crypt depth was 35 microns. Both these values are considerably less than those found in adult intestine. These results provide strong support for the general hypothesis that some factor associated with the physical length of the crypt, called crypt factor or CF, is directly responsible for controlling the way enterocytes organize subsequent structural differentiation of their surface membranes.

Aging↗

Ecological and social effects on reproduction and local recruitment in the red-backed shrike.

Numerous hypotheses have been proposed to explain variation in reproductive performance and local recruitment of animals. While most studies have examined the influence of one or a few social and ecological factors on fitness traits, comprehensive analyses jointly testing the relative importance of each of many factors are rare. We investigated how a multitude of environmental and social conditions simultaneously affected reproductive performance and local recruitment of the red-backed shrike Lanius collurio (L.). Specifically, we tested hypotheses relating to timing of breeding, parental quality, nest predation, nest site selection, territory quality, intraspecific density and weather. Using model selection procedures, predictions of each hypothesis were first analysed separately, before a full model was constructed including variables selected in the single-hypothesis tests. From 1988 to 1992, 50% of 332 first clutches produced at least one fledgling, while 38.7% of 111 replacement clutches were successful. Timing of breeding, nest site selection, predation pressure, territory quality and intraspecific density influenced nest success in the single-hypothesis tests. The full model revealed that nest success was negatively associated with laying date, intraspecific density, and year, while nest success increased with nest concealment. Number of fledglings per successful nest was only influenced by nest concealment: better-camouflaged nests produced more fledglings. Probability of local recruitment was related to timing of breeding, parental quality and territory quality in the single-hypothesis tests. The full models confirmed the important role of territory quality for recruitment probability. Our results suggest that reproductive performance, and particularly nest success, of the red-backed shrike is primarily affected by timing of breeding, nest site selection, and intraspecific density. This study highlights the importance of considering many factors at the same time, when trying to evaluate their relative contributions to fitness and life history evolution.

Altitude↗

Secondary endpoints cannot be validly analyzed if the primary endpoint does not demonstrate clear statistical significance.

There is lack of consensus surrounding the interpretation of observed treatment effects for secondary clinical endpoints when the primary endpoint for which the clinical trial was initially designed does not meet the objective of a demonstrated effect. We provide some arguments to support caution in making inferences for secondary endpoints in this situation. We examine the definitions of primary and secondary endpoints within the context of a hypothesis-testing framework for multiple endpoints, and we address the relationship of the correlation structure of these endpoints and the statistical adjustments needed to preserve experiment-wise type I error for a valid inference. We also address the hypothesis-testing framework and the estimation framework for valid inference, focusing on the interpretation of p-values associated with differentially powered hypothesis tests for each endpoint to detect an important clinical effect. We point out the limitations on the strength of evidence (and quantification of uncertainty) for a secondary endpoint effect that can be derived from only one study and introduce the likelihood of replication of the finding in another study of identical size and design as a useful concept to guide this interpretation.

Clinical Trials as Topic↗

Explicit category learning in Parkinson's disease: deficits related to impaired rule generation and selection processes.

The present study examined the source of explicit category learning deficits previously noted in patients with Parkinson's disease (PD). Task stimuli consisted of 4 binary-valued cues that together determined category assignment, although some cues were more important for the categorization decision. Participants verbalized the hypotheses being tested to provide several measures of the hypothesis testing. Analyses of these verbal protocols indicated that PD patients were impaired on rule generation and selection but not rule shifting. Patients had particular difficulty noting the relative importance of the cues. Specific aspects of performance were differently correlated with neuropsychological measures of working memory and hypothesis testing ability. Together, the results suggest that the cognitive processes required for explicit category learning are mediated by partially distinct neural mechanisms.

Aged↗

Alcohol-abusing teenage boys. Testing a hypothesis on alcohol abuse and personality factors, using a personality inventory.

This is the second part of an investigation of alcohol-abusing teenage boys, focusing on personality. One group of 50 High-consumers and one group of 50 0-consumers were selected from 862 18-year-old boys in the general population summoned to the Regional Recruiting and Replacement Office in Solna. These boys answered a personality inventory (KSP) to test a hypothesis on alcohol abuse and personality factors which might indicate psychopathy. The results support the hypothesis that alcohol-abusing teenage boys have psychopathic personality traits while the non-consuming boys have normal personalities. The study cannot reveal whether the differences in personality were the result of the high alcohol consumption or if the psychopathic personality traits preceded the high consumption. A reasonable hypothesis for further research is that vulnerable boys living under poor social conditions react to their situation with motoric restlessness, impulsiveness and aggressive acting-out behaviour. Due to this their social adjustment as grown ups is poor with consequent alcohol and drug abuse and criminality.

Adolescent↗

Proceedings of the SMBE Tri-National Young Investigators' Workshop 2005. Control of the false discovery rate applied to the detection of positively selected amino acid sites.

In this article, we consider the probabilistic identification of amino acid positions that evolve under positive selection as a multiple hypothesis testing problem. The null hypothesis "H0,s: site s evolves under a negative selection or under a neutral process of evolution" is tested at each codon site of the alignment of homologous coding sequences. Standard hypothesis testing is based on the control of the expected proportion of falsely rejected null hypotheses or type-I error rate. As the number of tests increases, however, the power of an individual test may become unacceptably low. Recent advances in statistics have shown that the false discovery rate--in this case, the expected proportion of sites that do not evolve under positive selection among those that are estimated to evolve under this selection regime--is a quantity that can be controlled. Keeping the proportion of false positives low among the significant results generally leads to an increase in power. In this article, we show that controlling the false detection rate is relevant when searching for positively selected sites. We also compare this new approach to traditional methods using extensive simulations.

Amino Acids↗

Epidemiologic identification of occupational carcinogens.

Epidemiology has a role to play in the identification of occupational carcinogens through hypothesis testing and surveillance for new carcinogenic hazards. Hypothesis testing is undertaken mainly by retrospective (non-concurrent) cohort studies and case-control studies. The former are limited particularly by difficulties in follow-up and inadequacy of data on exposure to the agent of interest and possible confounding or interacting factors. The latter are limited mainly by the problem of bias in the retrospective determination of exposure. Several studies giving similar results are therefore usually required before an association can be considered with any confidence as established. Surveillance for new hazards may be maintained by the regular analysis of routinely collected cancer incidence or mortality data. For early detection of hazards this should be supplemented by special studies, either on-going case control studies of cancers which are commonly due to occupation (e.g. lung, bladder, liver and nasal cancers) or linkage of personnel records from high risk industries to cancer registry or death records.

Carcinogens, Environmental↗

An E-M algorithm and testing strategy for multiple-locus haplotypes.

This paper gives an expectation maximization (EM) algorithm to obtain allele frequencies, haplotype frequencies, and gametic disequilibrium coefficients for multiple-locus systems. It permits high polymorphism and null alleles at all loci. This approach effectively deals with the primary estimation problems associated with such systems; that is, there is not a one-to-one correspondence between phenotypic and genotypic categories, and sample sizes tend to be much smaller than the number of phenotypic categories. The EM method provides maximum-likelihood estimates and therefore allows hypothesis tests using likelihood ratio statistics that have chi 2 distributions with large sample sizes. We also suggest a data resampling approach to estimate test statistic sampling distributions. The resampling approach is more computer intensive, but it is applicable to all sample sizes. A strategy to test hypotheses about aggregate groups of gametic disequilibrium coefficients is recommended. This strategy minimizes the number of necessary hypothesis tests while at the same time describing the structure of disequilibrium. These methods are applied to three unlinked dinucleotide repeat loci in Navajo Indians and to three linked HLA loci in Gila River (Pima) Indians. The likelihood functions of both data sets are shown to be maximized by the EM estimates, and the testing strategy provides a useful description of the structure of gametic disequilibrium. Following these applications, a number of simulation experiments are performed to test how well the likelihood-ratio statistic distributions are approximated by chi 2 distributions. In most circumstances the chi 2 grossly underestimated the probability of type I errors. However, at times they also overestimated the type 1 error probability. Accordingly, we recommended hypothesis tests that use the resampling method.

Algorithms↗

Expertise in cognitive psychology: testing the hypothesis of long-term working memory in a study of soccer players.

This experiment compared several theories of expertise and exceptional performances in cognitive psychology. One current conception assumes that experts in a specific domain have developed a long-term working memory, which accounts for the difference in memory performance between experts and novices. The principal characteristics of this memory are the speed with which processes of storage and retrieval function and the existence of retrieval structures that allow a temporary activation of the knowledge store in long-term memory. Other authors such as Vicente and Wang argue this notion does not account for memory performance that is not intrinsic to the domain of expertise. We attempt to clarify the two viewpoints and to focus on this debate by testing the hypothesis of long-term working memory using soccer as the domain of expertise and by comparing the cognitive performance of participants who have different expertise (novices, supporters, players, and coaches). 35 male participants were administered a new version of the Reading Span test to assess their long-term working memory according to two conditions. In the first condition (structured condition), the last word of each sentence was related to the soccer domain, and these words were related to each other in such a manner that they represented a part of the game. In the second condition (unstructured condition), the last word of each sentence was related to soccer but these words did not represent part of the game. Analysis showed that the sentence span increased as a function of expertise for the structured condition but not for the unstructured condition. The results were interpreted in the framework of the constraint attunement hypothesis proposed by Vicente in 1992 and the long-term working memory hypothesis proposed by Ericsson and Kintsch in 1995.

Adult↗

Methodological issues in testing the hypothesis of risk compensation.

The hypothesis of risk compensation implies that persons experiencing a real or perceived change in the riskiness of an activity will alter their consumption of that activity to obtain a preferred combination of risk and reward. In evaluating whether individuals display compensating behavior in response to safety interventions, not all persons subject to the intervention will necessarily display compensating behavior, even if the hypothesis is correct: the hypothesis has testable implications only for the subset of persons subject to the intervention who perceive that their risk has changed. This paper argues that methodologies that include persons for whom the hypothesis has no testable implications (against a null hypothesis of no compensation effect) result in estimates of the compensation effect and test statistics which are biased towards zero. Previously published data on motor-vehicle-related injuries to cyclists and pedestrians in Britain before and after a mandatory safety-belt-use law went into effect were used to infer the size of this bias. In these data, the inclusion of persons for whom the hypothesis of risk compensation has no testable implications appears to have resulted in estimates of a risk-compensation effect which are too small by about half. This work suggests that the British data are consistent with a risk-compensation effect of 7-13 percent, and raises important methodological issues in testing the hypothesis of risk compensation.

Accidents, Traffic↗

Spearman's hypothesis and test score differences between whites, Indians, and blacks in South Africa.

Numerous studies in the United States have shown that mean test scores between Blacks and Whites differ by about one standard deviation. It has further been noted that the magnitudes of these differences vary on different tests. This variation can be explained by Spearman's hypothesis, which states that Black-White differences on a set of cognitive tests are positively associated with the tests' g loadings (the general intellectual ability). The present study, conducted among Black, Indian, and White secondary students in South Africa, showed mean Black-White differences of two standard deviations, indicating that the American results of one standard deviation are not universally correct. With regard to Spearman's hypothesis, it was found that, although the mean White-Indian differences were about one standard deviation, these differences did not support the hypothesis. Results pertaining to the Black-White differences were ambiguous; the correlation of .62 (p < .05) between the Black g and the Black-White differences strongly supported the hypothesis. A nonsignificant correlation of .23 was obtained between the White g and the Black-White differences. Possible reasons for this finding are discussed.

Adolescent↗

Exact unconditional inference for risk ratio in a correlated 2 x 2 table with structural zero.

In this article, we consider small-sample statistical inference for rate ratio (RR) in a correlated 2 x 2 table with a structural zero in one of the off-diagonal cells. Existing Wald's test statistic and logarithmic transformation test statistic will be adopted for this purpose. Hypothesis testing and confidence interval construction based on large-sample theory will be reviewed first. We then propose reliable small-sample exact unconditional procedures for hypothesis testing and confidence interval construction. We present empirical results to evince the better confidence interval performance of our proposed exact unconditional procedures over the traditional large-sample procedures in small-sample designs. Unlike the findings given in Lui (1998, Biometrics 54, 706-711), our empirical studies show that the existing asymptotic procedures may not attain a prespecified confidence level even in moderate sample-size designs (e.g., n = 50). Our exact unconditional procedures on the other hand do not suffer from this problem. Hence, the asymptotic procedures should be applied with caution. We propose two approximate unconditional confidence interval construction methods that outperform the existing asymptotic ones in terms of coverage probability and expected interval width. Also, we empirically demonstrate that the approximate unconditional tests are more powerful than their associated exact unconditional tests. A real data set from a two-step tuberculosis testing study is used to illustrate the methodologies.

Confidence Intervals↗

Predicting reproductive outcome from sperm measurements in Swiss (CD-1) mice.

UNLABELLED: Recent advances in the quantification of sperm characteristics, particularly by techniques of videomicrography, raise the problem of identifying a subset of sperm measurements that can accurately predict reduced reproductive performance in the presence of a reproductive toxicant. This paper discusses and illustrates, with sperm data from Swiss (CD-1) mice, three properties of sperm measurements in addition to the association with reproductive outcome that can be used as objective criteria to distinguish among potentially useful sperm characteristics. Identification of these properties was motivated by the need to single out sperm characteristics that would yield hypothesis tests with good power to distinguish groups with altered sperm characteristics from those with normal sperm characteristics. The list of desirable properties of sperm characteristics includes: a) MEASUREMENT: The most useful sperm characteristics will exhibit low measurement bias and high measurement precision. b) Distribution: Sperm characteristics that follow the class of normal distributions allow easy identification of the most powerful hypothesis testing procedures. c) Variation: Sperm characteristics that exhibit limited variation from individual to individual will make altered values easier to detect. d) Correlation: Sperm characteristics that exhibit a high correlation with reproductive outcome will be most useful. MEASUREMENTs of sperm concentration, sperm motility, and abnormal sperm were examined for each of these properties in Swiss (CD-1) mice. For control animals, evidence of measurement bias between labs and significant variation among studies within labs was found for each sperm characteristic, with measurements of abnormal sperm exhibiting the least bias. Sperm concentration and the natural logarithm of abnormal sperm appeared normally distributed. Individual to individual variation was substantial for sperm motility and sperm concentration, both of which required an approximately 30% decrease in mean to achieve good statistical power. In contrast, abnormal sperm required only a 4% increase in mean to achieve the same power. In spite of the measurement noise associated with these sperm characteristics, data from 25 experiments indicated good agreement between the results of hypothesis tests based on sperm characteristics and reproductive outcome as judged by fertility and the number of pups. The same conclusions reached for reproductive outcome were reached in 19 of 24 experiments for sperm motility, in 19 of 25 experiments for sperm concentration, and in 17 of 24 experiments for abnormal sperm.

Animals↗