Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

A robust sequential test for text-independent speaker verification.

A robust speaker verification algorithm based on sequential hypothesis testing is presented. In speaker verification, the system performance is severely degraded by deviations from the nominal statistical speaker models caused by insufficient training data, varying microphone and transmission line characteristics, and different levels and types of background noise. Historically, this problem has been addressed by using an empirical method to determine an adequate decision threshold for a desired operating point. The robust detection algorithm presented here is based on a minimax criterion, such that the worst-case performance over a class of distributions is minimized. The sequential test provides further robustness by using additional data if a decision cannot be reached with the desired level of confidence. The sequential detector has been implemented and tested on a realistic database collected specifically for this purpose, and the performance has been shown to be comparable to or better than that of the corresponding heuristic detectors previously described in the literature.

Algorithms↗

Confidence intervals and test of hypotheses concerning dose response relations inferred from animal carcinogenicity data.

Confidence intervals and hypothesis tests are developed for dose-response relations based on dichotomous data from animal carcinogenicity experiments. The functional form of the dose-response curve comes from the Armitage-Doll multistage carcinogenesis model and involves a polynomial in the dose-rate, with non-negative coefficients. Asymptotic distributions of the maximum likelihood estimators of these coefficients are used to construct confidence bounds on risk at a given dose and on the dose corresponding to a given risk. Likelihood ratio tests are developed for the presence of a positive dose-related effect and for the existence of a positive slope to the dose-response curve at zero dose. The latter test is of practical importance since a positive slope of the dose-response curve at zero dose rules out any "threshold-like" behavior and would often mean that any concentration low enough to insure a negligibly low cancer risk (e.g., 10(-6)) would be too low to be economically useful for applications such as food additives. Simulation experiments are performed to provide guidelines for applying the theory.

Animals↗

Most powerful permutation invariant tests for relatedness hypotheses using genotypic data.

The problem of inferring kinship structure among a sample of individuals using genetic markers is considered with the objective of developing hypothesis tests for genetic relatedness with nearly optimal properties. The class of tests considered are those that are constrained to be permutation invariant, which in this context defines tests whose properties do not depend on the labeling of the individuals. This is appropriate when all individuals are to be treated identically from a statistical point of view. The approach taken is to derive tests that are probably most powerful for a permutation invariant alternative hypothesis that is, in some sense, close to a null hypothesis of mutual independence. This is analagous to the locally most powerful test commonly used in parametric inference. Although the resulting test statistic is a U-statistic, normal approximation theory is found to be inapplicable because of high skewness. As an alternative it is found that a conditional procedure based on the most powerful test statistic can calculate accurate significance levels without much loss in power. Examples are given in which this type of test proves to be more powerful than a number of alternatives considered in the literature, including Queller and Goodknight's (1989) estimate of genetic relatedness, the average number of shared alleles (Blouin, 1996), and the number of feasible sibling triples (Almudevar and Field, 1999).

Alleles↗

When to combine hypotheses and adjust for multiple tests.

OBJECTIVE: To provide guidelines for identifying composite hypotheses and addressing the probability of false rejection for multiple hypotheses. DATA SOURCES AND STUDY SETTING: Examples from the literature in health services research are used to motivate the discussion of composite hypothesis tests and multiple hypotheses. METHODS: This article is a didactic presentation. PRINCIPAL FINDINGS: It is not rare to find mistaken inferences in health services research because of inattention to appropriate hypothesis generation and multiple hypotheses testing. Guidelines are presented to help researchers identify composite hypotheses and set significance levels to account for multiple tests. CONCLUSIONS: It is important for the quality of scholarship that inferences are valid: properly identifying composite hypotheses and accounting for multiple tests provides some assurance in this regard.

Bias↗

Anxiety and sport performance.

From the findings summarized in this review, it appears that there is little evidence in support of the inverted-U hypothesis. Available research indicates that there is considerable variability in the optimal precompetition anxiety responses among athletes, which does not conform to the inverted-U hypothesis. Many athletes appear to perform best when experiencing high levels of anxiety and interventions that act to produce quiescence may actually worsen the performance of this group. These findings indicate that there is a need to shift the research paradigm away from theories of anxiety and performance based on task characteristics or group effects and, instead, employ theoretical models that account for individual differences. Hanin's [39, 40] ZOF theory appears to be a good candidate for furthering our knowledge in this area. It was developed on the basis of research with athletes and it explicitly incorporates the concept of individual differences in the anxiety-performance relationship. Most important, because an individual's optimal range of anxiety is precisely defined, the validity of ZOF theory can be directly examined through hypothesis testing, whereas it has been argued that the inverted-U hypothesis is effectively shielded against falsification [84]. Although the findings of ZOF theory indicate that a significant percentage of athletes perform best at high levels of anxiety, Hanin's translated writings do not provide an explanation of why this is so. Further research is clearly indicated, but one explanation for this finding may involve how the athlete interprets or conceptualizes anxiety. For example, Mahoney and Avener [64] found that, although the absolute level of precompetition anxiety was similar between successful and unsuccessful Olympic gymnasts, there were differences in the way the athletes conceptualized the anxiety they were experiencing. The better performers viewed their anxiety as desirable, whereas anxiety was associated with self-doubts and catastrophizing in the unsuccessful gymnasts. Similar differences have been observed in the test anxiety literature where it has been found that poorer test takers perceive their anxiety to be more threatening and debilitating than do better performers [45]. Furthermore, temporal differences in the patterning of anxiety [64], fear responses, or cardiorespiratory measures [28] have been found between successful and unsuccessful performers; this may reflect a difference in the ability to regulate anxiety. It may also be the case that performance is not so much affected by the absolute level of precompetition anxiety as the consistency in the anxiety level across competitions. Athletes may also develop coping strategies that exploit consistent changes in attentional focus that result from elevated anxiety.(ABSTRACT TRUNCATED AT 400 WORDS)

Adult↗

On the non-inferiority of a diagnostic test based on paired observations.

Non-inferiority of a diagnostic test to the standard or the optimum test is a common issue in medical research. Often we want to determine if a new diagnostic test is as good as the standard reference test. Sometimes we are interested in an inexpensive test that may have an acceptably inferior sensitivity or specificity. While hypothesis testing procedures and sample size formulae for the equivalence of sensitivity or specificity alone have been proposed, very few studies have discussed simultaneous comparisons of both parameters. In this paper, we present three different testing procedures and sample size formulae for simultaneous comparison of sensitivity and specificity based on paired observations and with known disease status. These statistical procedures are then used to compare two classification rules that identify women for future osteoporotic fracture. Simulation experiments demonstrate that the new tests and sample size formulae give the appropriate type I and II error rates. Differences between our approach and the approach of Lui and Cumberland are discussed.

Clinical Trials as Topic↗

Comparison of probability and likelihood models for peptide identification from tandem mass spectrometry data.

We evaluate statistical models used in two-hypothesis tests for identifying peptides from tandem mass spectrometry data. The null hypothesis H(0), that a peptide matches a spectrum by chance, requires information on the probability of by-chance matches between peptide fragments and peaks in the spectrum. Likewise, the alternate hypothesis H(A), that the spectrum is due to a particular peptide, requires probabilities that the peptide fragments would indeed be observed if it was the causative agent. We compare models for these probabilities by determining the identification rates produced by the models using an independent data set. The initial models use different probabilities depending on fragment ion type, but uniform probabilities for each ion type across all of the labile bonds along the backbone. More sophisticated models for probabilities under both H(A) and H(0) are introduced that do not assume uniform probabilities for each ion type. In addition, the performance of these models using a standard likelihood model is compared to an information theory approach derived from the likelihood model. Also, a simple but effective model for incorporating peak intensities is described. Finally, a support-vector machine is used to discriminate between correct and incorrect identifications based on multiple characteristics of the scoring functions. The results are shown to reduce the misidentification rate significantly when compared to a benchmark cross-correlation based approach.

Databases, Protein↗

Associations of race, education, and patterns of preventive service use with stage of cancer at time of diagnosis.

OBJECTIVE: To go beyond the documentation of disparities by race and SES by analyzing health behaviors regarding preventive and cancer screening services and determining if these behaviors are associated with stage of cancer when first diagnosed. DATA: Stage of cancer for Medicare patients diagnosed in 1995 with breast, colorectal, uterine, ovarian, prostate, bladder, or stomach cancer; and use of influenza and pneumonia immunization, mammography, pap smear, colon cancer screening, and the prostate specific antigen test during the two years preceding diagnosis of cancer. STUDY DESIGN: Hypothesis tested: health behaviors regarding use of preventive and cancer screening services are associated with stage of cancer when first diagnosed. DATA COLLECTION/EXTRACTION METHODS: Information was extracted from the database formed by the linkage of Surveillance, Epidemiology, and End Results (SEER) cancer registries with Medicare files. PRINCIPAL FINDINGS: Black and white patients (of higher and lower SES) who used more of the preventive and cancer screening services were at a lower risk of having late stage cancer for six cancers studied (breast, colorectal [male and female], prostate, uterine, and male bladder cancer) than their counterparts who used fewer of these services. CONCLUSIONS: The use of preventive and cancer screening services is a health behavior associated with better health outcomes for the elderly diagnosed with cancer. The lack of preventive service use can serve as a marker for identifying persons at risk of late stage cancer when first diagnosed. Strategies that encourage the use of preventive services by low users of these services are likely to reinforce a range of healthy behaviors that help to ameliorate disparities in health outcomes.

Black or African American↗

Insulin-like growth factor binding protein 2: gene expression microarrays and the hypothesis-generation paradigm.

A major goal of modern medicine is to identify key genes and their products that are altered in the diseased state and to elucidate the molecular mechanisms underlying disease development, progression, and resistance to therapy. This is a daunting task given the exceptionally high complexity of the human genome. The paradigm for research has historically been hypothesis-driven despite the fact that the hypotheses under scrutiny often rest on tenuous subjective grounds or are derived from and dependent on chance observation. The imminent deciphering of the complete human genome, coupled with recent advances in high-throughput bioanalytical technology, has made possible a new paradigm in which data-based hypothesis-generation is the initial step in the investigative process, followed by hypothesis-testing. Genomics technologies are the primary source of the new hypothesis-generating capabilities that are now empowering biomedical researchers. The synergistic interaction between contemporary genomics technologies and the hypothesis-generation paradigm is well-illustrated by the discovery and subsequent ongoing study of the role of insulin-like growth factor binding protein 2 (IGFBP2) in human glioma biology. Using gene expression microarray technology, the IGFBP2 gene was recently found to be highly and differentially overexpressed in the most advanced grade of human glioma, glioblastoma. Based on this discovery, subsequent functional studies were initiated that suggest that IGFBP2 overexpression may contribute to the invasive nature of glioblastoma, and that IGFBP2 may exert its function via a newly identified novel binding protein. The IGFBP2 story is but one example of the power and potential of the new molecular methodologies that are transforming modern diagnostic and investigative neuropathology.

Brain Neoplasms↗

Effects of cytokines on antiviral pharmacokinetics: an alternative approach to assessment of drug interactions using bioequivalence guidelines.

The effects of cytokines on the pharmacokinetics of nucleoside analogs were evaluated in two separate studies using zidovudine in combination with interleukin-2 and didanosine in combination with alpha interferon. In each study, drug interactions were evaluated by using both a standard method (Student's t test) and bioequivalence testing. Serial blood samples were collected from human immunodeficiency virus-infected patients prior to and during cytokine therapy for determination of nucleoside analog concentrations. Concentrations were fit separately to a two-compartment model by using the iterative two-stage approach to population analysis. No alterations in area under the curve or oral clearance were observed for either drug during combination therapy. In general, there was good agreement between statistical methods for determining if antiviral pharmacokinetic parameters were altered by concomitant cytokine therapy. However, large individual changes in the maximum concentration of zidovudine in serum were detected by bioequivalence testing but no difference was found by Student's t test. For didanosine, significant but clinically irrelevant decreases determined by standard hypothesis testing were seen for both the volume of the central compartment (1.91 to 1.86 liters) and the absorption rate constant (0.79 to 0.73 h-1) in the presence of alpha interferon. No interaction was noted for these parameters by using bioequivalence guidelines. Bioequivalence testing may provide an alternative approach to assessment of drug interactions. Interleukin-2 and alpha interferon do not alter the pharmacokinetics of zidovudine and didanosine, respectively.

Adult↗

Statistical analysis of categorical data.

This article has presented the most widely used techniques for analyzing categorical data. Data that are qualitative are naturally categorical. Even continuous (or measurement) data can be classified into categories (eg, ages can be grouped into above and below 40) and, hence, these procedures can be applied to a wide range of data. All of these statistical tests are classified as nonparametric statistics, although many statistical textbooks treat hypothesis testing of categorical data separately from the more common nonparametric methods. The most widely used nonparametric procedures will be reviewed in the next issue.

Data Interpretation, Statistical↗

Electrical stimulation of rhesus monkey nucleus reticularis gigantocellularis. II. Effects on metrics and kinematics of ongoing gaze shifts to visual targets.

Saccade kinematics are altered by ongoing head movements. The hypothesis that a head movement command signal, proportional to head velocity, transiently reduces the gain of the saccadic burst generator (Freedman 2001, Biol Cybern 84:453-462) can account for this observation. Using electrical stimulation of the rhesus monkey nucleus reticularis gigantocellularis (NRG) to alter the head contribution to ongoing gaze shifts, two critical predictions of this gaze control hypothesis were tested. First, this hypothesis predicts that activation of the head command pathway will cause a transient reduction in the gain of the saccadic burst generator. This should alter saccade kinematics by initially reducing velocity without altering saccade amplitude. Second, because this hypothesis does not assume that gaze amplitude is controlled via feedback, the added head contribution (produced by NRG stimulation on the side ipsilateral to the direction of an ongoing gaze shift) should lead to hypermetric gaze shifts. At every stimulation site tested, saccade kinematics were systematically altered in a way that was consistent with transient reduction of the gain of the saccadic burst generator. In addition, gaze shifts produced during NRG stimulation were hypermetric compared with control movements. For example, when targets were briefly flashed 30 degrees from an initial fixation location, gaze shifts during NRG stimulation were on average 140% larger than control movements. These data are consistent with the predictions of the tested hypothesis, and may be problematic for gaze control models that rely on feedback control of gaze amplitude, as well as for models that do not posit an interaction between head commands and the saccade burst generator.

Animals↗

The randomization test applied to flow cytometric histograms.

The randomization test is used to test the hypothesis that two groups of flow cytometric histograms consist of samples from the same probability density function. The hypothesis is tested channel-by-channel. This non-parametric method does not require any assumptions as to the probability density functions involved and is therefore applicable to histograms from many different biological systems. Moreover, it allows for very few histograms per group so that the hypothesis can be tested on the basis of a small number of experiments. The procedure is implemented in PASCAL on a mini computer which is connected to a flow cytometer.

Computers↗

Evaluation of hormonal testing in the screening for in vitro fertilization (IVF) of women with tubal factor infertility.

PURPOSE: To evaluate the frequency of abnormal prolactin and thyroid stimulating hormone (TSH) test results in ovulatory women with tubal factor infertility who were screened for in vitro fertilization (IVF). METHODS: Charts were identified from 112 ovulatory women with follicle stimulating hormone (FSH) < 20 mIU/ml who were diagnosed with tubal factor infertility and were screened for IVF with thyroid stimulating hormone (TSH) and prolactin levels. Women previously diagnosed with thyroid disease were subsequently excluded and 98 subjects remained. All subjects were determined to be ovulatory by biphasic basal body temperature (BBT) charts, luteal phase progesterone > 4 ng/ml, or endometrial biopsy revealing secretory endometrium. Results of cycle day 3 serum TSH and prolactin concentrations were recorded. The normal range for each test reflects the geometric mean +/- 2 standard deviations (i.e., 95% interval), as obtained from the reference laboratory. Under this construct, hypothesis tests were performed to determine whether or not our study population was consistent with the reference range of normal hormone levels. Under the null hypothesis (normal levels), we expected 5% of the TSH tests to be abnormal (i.e., high or low levels), and 2.5% of the prolactin tests to be abnormal (i.e., high levels). Exact bionomial confidence intervals and P-values were calculated. We also tested for age trend in the proportion of abnormal results. RESULTS: Study subjects had an age range of 25-43. In the study group, 4 (0.041) out of the 98 women screened had abnormal TSH levels. Of these four abnormal TSH results, three were elevated (i.e., TSH > 4.6 microIU/ml) and one was low (i.e., TSH < 0.6 microIU/ml). The frequency of an abnormal TSH value was not significantly different from that expected from the reference laboratory normal values (95% CI 0.011, 0.101). Of the 98 subjects, 7 (0.071) had abnormal prolactin levels, which was significantly different from that expected from the reference laboratory normal values (95% CI 0.029, 0.142; P = 0.023). When stratified by age, there was no observed trend of abnormalities for TSH or prolactin levels with increasing age. CONCLUSIONS: In ovulatory women presenting for IVF with tubal factor infertility, our results show that routine screening with a TSH test does not yield a significantly higher proportion of abnormal results than that expected from the reference laboratory normal values. However, prolactin level screening was found to yield a higher incidence of abnormal tests than expected from the reference laboratory normal values.

Adult↗

Are tumor-associated transplantation antigens of chemically induced sarcomas related to alien histocompatibility antigens?

The hypothesis tested was that tumor-specific transplantation antigens of chemically induced tumors cross-react with allogeneic histocompatibility antigens. This hypothesis makes several predictions that can be tested experimentally. First, tumors should grow better and be less immunogenic in certain F1 hybrids than in their syngeneic parents, owing to the hypothecated cross-reactivity of the tumor-specific transplantation antigens with F1 antigens. This is in contrast to the more common observation that parental strain tumors grow worse in the F1 hybrids than they do in the parent. Also certain allogeneic skin grafts might immunize the parental strain mice against their syngeneic tumors, and, finally, immunizing parental mice with syngeneic tumor might cause accelerated rejection of certain skin allografts. The results show that certain tumors grew better in the F1 mice than they did in the parents but that the tumors were not less immunogenic in the F1 hybrids. Mice immunized against alloantigens showed a dose-dependent enhancement of syngeneic tumor growth. Finally, mice immunized with syngeneic tumors demonstrated an apparent prolongation of certain skin allografts. The discussion considers possible alternatives explaining these results.

Animals↗

The geography of power: statistical performance of tests of clusters and clustering in heterogeneous populations.

Heterogeneous population densities complicate comparisons of statistical power between hypothesis tests evaluating spatial clusters or clustering of disease. Specifically, the location of a cluster within a heterogeneously distributed population at risk impacts power properties, complicating comparisons of tests, and allowing one to map spatial variations in statistical power for different tests. Such maps provide insight into the overall power of a particular test, and also indicate areas within the study area where tests are more or less likely to detect the same local increase in relative risk. While such maps are largely driven by local sample size, we also find differences due to features of the statistics themselves. We illustrate these concepts using two tests: Tango's index of clustering and the spatial scan statistic. Furthermore, assessments of the accuracy of the 'most likely cluster' involve not only statistical power, but also spatial accuracy in identifying the location of a true underlying cluster. We illustrate these concepts via induction of artificial clusters within the observed incidence of severe cardiac birth defects in Santa Clara County, CA in 1981.

California↗

Familial analysis of qualitative traits under multifactorial inheritance.

The analysis of family data is described for qualitative multifactorial traits. For such a trait, affectational status is determined by an underlying liability distribution with one or more thresholds. The distribution of families (either selected at random or through probands) is used to estimate sex-specific parent-offspring and sibling correlations in liability and the prevalence in each sex. In contrast to using pairs of relatives, this approach permits estimation of age-specific population prevalences without a control sample. Moreover, by allowing for sex-specific correlations and a correlation between mates, path analysis can be used to model and test various cultural transmission models in addition to polygenic inheritance. Parameter estimation, hypothesis testing, and a goodness-of-fit test for path analytic models are described, and a computer program implementing these procedures is outlined.

Age Factors↗

Meehl on metatheory.

This article proffers a critical survey of Paul E. Meehl's thoughts on the philosophy of science, as well as constructive criticisms of his views on corroboration of scientific theories. Using examples from clinical psychology and allied domains, six major topics are addressed: (a) the nature of theoretical constructs, (b) statistical diagnosis of clinical categories, (c) detection of hypothesized taxa, (d) null-hypothesis significance testing, (e) complexified hypothesis testing, and (f) cliometric theory appraisal. Through numerous examples, it is shown that Paul Meehl had a superb ability to recognize, articulate, and clarify hidden complexities, and other underappreciated obscurities in important working concepts.

Forecasting↗