Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Screening for possible human carcinogens and mutagens. False positives, false negatives: statistical implications.

A screening method aimed at identifying potential human carcinogens using either animal cancer bioassays or short-term genotoxic assays has 4 possible results: true positive, true negative, false positive and false negative. Such a categorisation is superficially similar to the results of hypothesis testing in a statistical analysis. In this latter case the false positive rate is determined by the significance level of the test and the false negative rate by the statistical power of the test. Although the two types of categorisation appear somewhat similar, different statistical issues are involved in their interpretation. Statistical methods appropriate for the analysis of the results of a series of assays include the use of Bayes' theorem and multivariate methods such as clustering techniques for the selection of batteries of short-term test capable of a better prediction of potential carcinogens. The conclusions drawn from such studies are dependent upon the estimates of values of sensitivity and specificity used, the choice of statistical method and the nature of the data set. The statistical issues resulting from the analysis of specific genotoxicity experiments involve the choice of suitable experimental designs and appropriate analyses together with the relationship of statistical significance to biological importance. The purpose of statistical analysis should increasingly be to estimate and explore effects rather than for formal hypothesis testing.

Bayes Theorem↗

Robust asymptotic sampling theory for correlations in pedigrees.

Methods to unravel the genetic determinants of non-Mendelian diseases lie at the next frontier of statistical approaches for human genetics. It is generally agreed that, before proceeding with segregation or linkage analysis, the trait under study ought to be shown to exhibit familial correlation. By coding dichotomous traits as binary variables, a single robust approach in the estimation of pedigree correlations, rather than two distinct approaches, can be used to assess the potential heritability of a trait, and, latterly, to examine the mode of inheritance. The asymptotic theory to conduct hypothesis tests and confidence intervals for correlations among different members of nuclear families is well established but is applicable only if the nuclear families are independent. As a further contribution to the literature, we derive the asymptotic sampling distribution of correlations between random variables among arbitrary pairs of members in extended families for the Pearson product-moment estimator with generalized weights. This derivation is done without assuming normality of the traits. The sampling distribution is shown to be asymptotically normal to first order, and hence large-sample hypothesis tests and confidence intervals with estimates of the variances and correlation coefficients are proposed. Discussion concludes with an example and a suggestion for future research.

ABO Blood-Group System↗

A comprehensive literature review of haplotyping software and methods for use with unrelated individuals.

Interest in the assignment and frequency analysis of haplotypes in samples of unrelated individuals has increased immeasurably as a result of the emphasis placed on haplotype analyses by, for example, the International HapMap Project and related initiatives. Although there are many available computer programs for haplotype analysis applicable to samples of unrelated individuals, many of these programs have limitations and/or very specific uses. In this paper, the key features of available haplotype analysis software for use with unrelated individuals, as well as pooled DNA samples from unrelated individuals, are summarised. Programs for haplotype analysis were identified through keyword searches on PUBMED and various internet search engines, a review of citations from retrieved papers and personal communications, up to June 2004. Priority was given to functioning computer programs, rather than theoretical models and methods. The available software was considered in light of a number of factors: the algorithm(s) used, algorithm accuracy, assumptions, the accommodation of genotyping error, implementation of hypothesis testing, handling of missing data, software characteristics and web-based implementations. Review papers comparing specific methods and programs are also summarised. Forty-six haplotyping programs were identified and reviewed. The programs were divided into two groups: those designed for individual genotype data (a total of 43 programs) and those designed for use with pooled DNA samples (a total of three programs). The accuracy of programs using various criteria are assessed and the programs are categorised and discussed in light of: algorithm and method, accuracy, assumptions, genotyping error, hypothesis testing, missing data, software characteristics and web implementation. Many available programs have limitations (eg some cannot accommodate missing data) and/or are designed with specific tasks in mind (eg estimating haplotype frequencies rather than assigning most likely haplotypes to individuals). It is concluded that the selection of an appropriate haplotyping program for analysis purposes should be guided by what is known about the accuracy of estimation, as well as by the limitations and assumptions built into a program.

Algorithms↗

Estimation and detection of event-related fMRI signals with temporally correlated noise: a statistically efficient and unbiased approach.

Recent developments in analysis methods for event-related functional magnetic resonance imaging (fMRI) has enabled a wide range of novel experimental designs. As with selective averaging methods used in event-related potential (ERP) research, these methods allow for the estimation of the average time-locked response to particular event-types, even when these events occur in rapid succession and in an arbitrary sequence. Here we present a flexible framework for obtaining efficient and unbiased estimates of event-related hemodynamic responses, in the presence of realistic temporally correlated (nonwhite) noise. We further present statistical inference methods based upon the estimated responses, using restriction matrices to formulate temporal hypothesis tests about the shape of the evoked responses. The accuracy of the methods is assessed using synthetic noise, actual fMRI noise, and synthetic activation in actual noise. Actual false-positive rates were compared to nominal false-positive rates assuming white noise, as well as local and global noise estimates in the estimation procedure (assuming white noise resulted in inappropriate inference, while both global and local estimates corrected false-positive rates). Furthermore, both local and global noise estimates were found to increase the statistical power of the hypothesis tests, as measured by the receiver operating characteristics (ROC). This approach thus enables appropriate univariate statistical inference with improved statistical power, without requiring a priori assumptions about the shape or timing of the event-related hemodynamic response.

Artifacts↗

What is confidence? Part 1: The use and interpretation of confidence intervals.

Hypothesis testing and the P value it generates are overemphasized in statistical analyses published in medical journals. An alternative, the confidence interval (CI), offers significantly more information to readers interpreting results. There have been many authoritative calls for the report of CIs in place of P values, such as that of the International Committee of Medical Journal Editors, whose guidelines for statistical reporting give the following instructions: "When possible, quantify findings and present them with appropriate indicators of measurement error or uncertainty (such as confidence intervals)," and "Avoid sole reliance on statistical hypothesis testing, such as the use of P values, which fails to convey important quantitative information." In this article, part 1, we provide an overview of CIs for the clinician reading the medical literature. We describe the advantages of CIs and explain and illustrate their proper interpretation. In part 2, which follows this article, we provide added information important for clinical researchers, including a precise definition of CIs, a compact reference to methods for calculating CIs in common situations, and an explanation of the difference between CIs and probability intervals.

Biometry↗

"Objective" methods and "subjective" experiences.

Current psychiatric research and practice emphasize measurement of operationalized variables, quantification, and rigorous hypothesis testing. The fact that mental states can be subjectively experienced and that thoughts can refer to things and events outside the mind suggests that such objectifying methods alone may not provide a complete approach to mental life. Other complementary but systematic methods can be described which stress that (1) words are often natural expressions, not labels, of experiences; (2) usefulness, not agreement with observation, can sometimes validate psychological expressions; (3) some data can only be gathered by interactive involvement, not dispassionate observation; (4) a goal of inquiry can be interpretation, not hypothesis testing; and (5) understanding may require a holistic approach which expands rather than constricts the realm of relevant data.

Adaptation, Psychological↗

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem↗

Strength of evidence for density dependence in abundance time series of 1198 species.

Population limitation is a fundamental tenet of ecology, but the relative roles of exogenous and endogenous mechanisms remain unquantified for most species. Here we used multi-model inference (MMI), a form of model averaging, based on information theory (Akaike's Information Criterion) to evaluate the relative strength of evidence for density-dependent and density-independent population dynamical models in long-term abundance time series of 1198 species. We also compared the MMI results to more classic methods for detecting density dependence: Neyman-Pearson hypothesis-testing and best-model selection using the Bayesian Information Criterion or cross-validation. Using MMI on our large database, we show that density dependence is a pervasive feature of population dynamics (median MMI support for density dependence = 74.7-92.2%), and that this holds across widely different taxa. The weight of evidence for density dependence varied among species but increased consistently, with the number of generations monitored. Best-model selection methods yielded similar results to MMI (a density-dependent model was favored in 66.2-93.9% of species time series), while the hypothesis-testing methods detected density dependence less frequently (32.6-49.8%). There were no obvious differences in the prevalence of density dependence across major taxonomic groups under any of the statistical methods used. These results underscore the value of using multiple modes of analysis to quantify the relative empirical support for a set of working hypotheses that encompass a range of realistic population dynamical behaviors.

Ecosystem↗

Statistical inference on mean dioptric power: asymmetric powers and singular covariance.

Methods have been developed recently for testing hypotheses on mean dioptric power and for constructing confidence regions in situations that are most likely to be encountered. In this paper the methods are extended to make the analysis complete. A new situation covered specifically is that of dioptric power not of the form sphere/cylinder x axis. Such powers, termed asymmetric powers because the dioptric power matrices are asymmetric, include the equivalent power of a thick obliquely crossed bitoric lens. A second situation is that in which the covariance matrix of the sample of powers is singular. Symmetric dioptric power (the more familiar form of power) can be represented by a point in three-dimensional space. In general, however, dioptric power is four dimensional in character. Singularity of covariance arises when variation in the sample is limited to a subspace of dimension less than the full three or four. The space spanned by the sample is called the range space of the sample. The dimension of the range space may be four, three, two, one or zero. Each case is considered in turn. Numerical examples of hypothesis testing are presented in range spaces of dimension four to one. The test statistic devised for each case also gives the equation of the confidence region about the mean of a sample of dioptric powers. Singularity can sometimes be avoided merely by taking larger samples and by taking more accurate readings. The problem of near singularity is briefly discussed. The paper allows basic hypothesis testing on mean dioptric power and the construction of confidence regions in all possible circumstances.

Multivariate Analysis↗

A multivariate method for measurement error correction using pairs of concentration biomarkers.

PURPOSE: Measurement error is a pervasive problem in behavioral epidemiology, and available methods of correction all have generally untenable assumptions. We propose a multivariate method with more realistic assumptions. METHODS: The method uses two concentration biomarkers for each nutritional variable of interest and structural equation modeling. This produces corrected estimates of the effects on an outcome variable of changing the true exposure variables by one standard deviation, a standardized regression calibration. However, hypothesis testing in original units is preserved. The main assumptions are that certain error correlations between dietary estimates and biomarkers or between biomarkers be close to zero. RESULTS: Two illustrative models used simulated data with the covariance structure of a real data set. The corrections produced often were very substantial. A sensitivity analysis allowed error correlations to depart from zero over a modest range. Root mean square biases show the advantage of the corrected approach. Relatively large calibration studies are needed for adequate precision. CONCLUSIONS: As long as concentration biomarkers are selected carefully, error-corrected multivariate hypothesis testing and standardized effect estimation is possible. With the deviations from assumptions that were tested, the corrected method usually produces much less biased results than an uncorrected analysis.

Bias↗

Variance component modeling of attachment level measurements.

The purpose of this study was to investigate the influence of within subject and within tooth variability of attachment-level measurements on hypothesis testing. Full-mouth attachment-level measurements were obtained at 4 sites per tooth in early onset periodontitis subjects (both localized N = 89 and generalized N = 139) and adult periodontitis subjects (N = 309). Variance component models utilizing iterative generalized least squares were employed to estimate the % of variance distributed at the level of the site, tooth and subject in the 3 subject populations. The data indicate that a significant % of the variance is distributed within the tooth and the subject. Ignoring the variance attributable to teeth or subjects can result in inappropriately low type I error rates for hypothesis testing. Thus, both the subjects and the tooth must be considered in the modeling of attachment level measurements. Also in studies in which a limited number of samples are taken, data analysis would be simplified if these samples were taken from different teeth.

Adult↗

[Statistical results: which method of presentation to chose?].

Hypothesis testing and significance is currently the most widely used method in the medical literature to report statistical results. However, this method has several limitations. The main one is linked to the risk of misinterpretation of the p value. The arbitrariness of the 5 percent value used to determine whether a result is or not statistically significant is not always kept in mind, and the concept of statistical significance might therefore be confused with that of clinical or biological relevance. The misinterpretation pitfalls are mostly linked to the fact that the p value does not give precise indications on the strength of the association and its direction, or on the variability in the sample. Therefore, some experts claim that hypothesis testing and significance should be avoided in reporting statistical results, and that the method based upon estimation and confidence interval should be more widely used. By this latter method, it is possible to know the direction of the association and the effect size (i.e. the strength of the association). The precision of the estimation, i.e. the variability of the estimation in the sample, can be assessed by the width of the confidence interval: the narrower the confidence interval, the more precise the estimation. Therefore, the clinical relevance of the findings is easier to infere from such results than from those only reporting p values. However, the estimation and confidence interval method is not without its own limitations. This method is difficult to apply to non-parametric tests, and for some results, such as the comparison of mortality ratios, the p value is highly informative. On the other hand, the misinterpretation risk is not totally ruled out when estimation and confidence interval method is used. In the situations where both methods can be employed, there is not yet in the scientific community a definite consensus on which method is the best one to report statistical results, hence some experts suggest that both methods can be presented simultaneously, especially for clinical and epidemiological studies.

Confidence Intervals↗

Synovial, chondropathic, depositional: the radiological categories of arthritis. A review.

The accepted optimum logic frame for complex diagnostic problems is sequential hypothesis testing. The X-ray diagnosis of arthritis has now become sufficiently complex to make this the procedure of choice, but standard classifications of arthritis lack the discriminatory power needed for its effective deployment. This obstacle can be overcome if these standard classifications are replaced by a more discriminant classification based on radiologically identifiable discriminators. This classification divides arthritis into three groups: synovial, chondropathic and depositional. The initial categorization is usually made fairly simply from two or three markers, of which the most important is the site of any erosions present. Once categorized, the subsequent analysis can proceed in an approximately binary manner using the discriminators appropriate for that category. The use of this radiological classification simplifies the diagnostic approach, reduces the workload, and provides the algorithm needed for sequential hypothesis testing.

Algorithms↗

A four-shock Bayesian up-down estimator of the 80% effective defibrillation dose.

INTRODUCTION: New defibrillation techniques are often compared to standard approaches using the defibrillation threshold. However, inference from thresholding data necessitates extrapolation from reactions to relatively ineffective shocks, an error prone procedure requiring large sample sizes for hypothesis testing and large safety margins for defibrillator implantation. In contrast, this article presents a clinically validated statistical model of a minimum error, four-shock defibrillation testing protocol for estimating the 80% effective defibrillation strength for a given patient (ED80). METHODS AND RESULTS: A Bayesian statistical model was constructed assuming that the defibrillation dose-response curve is sigmoidal, and the ED80 is between 150 and 750 V. The model was used to design a minimum predicted error testing protocol and estimates. To prospectively validate the testing protocol and estimates, 170 patients received voltage-programmed biphasic testing. Four fibrillation episodes were induced and terminated in each patient according to the Bayesian up-down protocol. In addition, a validation attempt was made at the estimated ED80 rounded up to the nearest 50 V. In order to estimate the safety margin, in 136 patients, a defibrillation attempt was made at the rounded ED80 + 100 V. Of the 170 attempts at the rounded ED80, 143 (84%) attempts terminated fibrillation. Of the 136 attempts at the rounded ED80 + 100 V, 133 (98%) were effective. CONCLUSIONS: The four-shock Bayesian up-down protocol is the first clinical protocol to accurately predict an ED80 voltage. A 100 V increment above the ED80 provides an adequate safety margin. This simple and accurate method for estimating a highly effective defibrillation dose may be a valuable tool for population-based clinical hypothesis testing, as well as defibrillator implantation.

Adult↗

Abolition of the receptor potential response of isolated mammalian outer hair cells by hair-bundle treatment with elastase: a test of the tip-link hypothesis.

To test the hypothesis that the tip-links of hair-cell stereocilia are essential for mechanoelectrical transduction, tip-links of isolated outer hair cells (OHCs) of the guinea-pig cochlea were eliminated with a proteolytic enzyme, elastase, and the influence on the receptor potential measured with the whole-cell patch-clamp technique. Within 45 s of immersion of the hair bundle in 20 IU/ml elastase, the receptor potential in response to direct deflection of the hair bundle was irreversibly abolished. The electrical input impedance of the cell remained unchanged, implying that the channels of the basolateral membrane were not affected by elastase. The effect of elastase on the receptor potential was comparable to changes seen after mechanically induced hair-bundle damage. As a further control, a putative transduction-channel blocker, dihydrostreptomycin (68 microM), which does not affect tip-links, was applied to the hair bundle. Although the receptor potential was also blocked by dihydrostreptomycin, the effect was reversible. The results suggest that tip-links are required for mechanoelectrical transduction of mammalian OHCs.

Animals↗

Overcoming feelings of powerlessness in "aging" researchers: a primer on statistical power in analysis of variance designs.

A general rationale and specific procedures for examining the statistical power characteristics of psychology-of-aging empirical studies are provided. First, 4 basic ingredients of statistical hypothesis testing are reviewed. Then, 2 measures of effect size are introduced (standardized mean differences and the proportion of variation accounted for by the effect of interest), and methods are given for estimating these measures from already-completed studies. Power and sample size formulas, examples, and discussion are provided for common comparison-of-means designs, including independent samples I-factor and factorial analysis of variance (ANOVA) design, analysis of covariance designs, repeated measures (correlated samples) ANOVA designs, and split-plot (combined between- and within-subjects) ANOVA designs. Because of past conceptual differences, special attention is given to the power associated with statistical interactions, and cautions about applying the various procedures are indicated. Illustrative power estimations also are applied to a published study from the literature. It is argued that psychology-of-aging researchers will be both better informed consumers of what they read and more "empowered" with respect to what they research by understanding the important roles played by power and sample size in statistical hypothesis testing.

Aged↗

Maintenance of meiotic arrest in mouse oocytes by purines: modulation of cAMP levels and cAMP phosphodiesterase activity.

Hypoxanthine and adenosine are present in preparations of mouse ovarian follicular fluid, and these purines maintain mouse oocytes in meiotic arrest in vitro (Eppig et al.: Biology of Reproduction 33:1041-1049. 1985). The first hypothesis tested in this study is that purines which maintain meiotic arrest act by maintaining meiosis-arresting levels of cyclic adenosine monophosphate (cAMP) in the oocyte. Oocyte-cumulus cell complexes were incubated in control medium (no added purines), or medium containing 0.75 mM adenosine, 4 mM hypoxanthine, or both for 3 hr and the percentage of the oocytes that underwent germinal vesicle breakdown (GVB) and the cAMP content of the intact complexes and the oocytes were determined. Adenosine alone had little inhibitory effect on GVB at this time point but sustained higher levels of cAMP in the oocytes. Hypoxanthine maintained 80% of cumulus cell-enclosed oocytes in meiotic arrest and also sustained higher cAMP levels in the oocytes. The addition of adenosine to hypoxanthine-containing medium increased the percentage of oocytes maintained in meiotic arrest, and increased the amount of cAMP in the oocytes above that maintained by either hypoxanthine or adenosine alone. Neither hypoxanthine, adenosine, nor hypoxanthine plus adenosine altered the cAMP content of intact complexes when assayed after 3 hr culture. Microinjection of an inhibitor of the catalytic subunit of cAMP-dependent protein kinase induced GVB in denuded oocytes cultured in medium containing hypoxanthine. This purine, therefore, maintained meiotic arrest by sustaining elevated cAMP levels within the oocytes. The second hypothesis tested in this study is that purines maintain meiosis-arresting levels of cAMP, at least in part, by inhibiting cAMP phosphodiesterase activity. In descending order of potency, 3-isobutyl-1-methylxanthine (IBMX), guanosine, hypoxanthine, adenosine, and xanthosine inhibited cAMP phosphodiesterase in oocyte lysates. Moreover, like the potent phosphodiesterase inhibitor IBMX, hypoxanthine augmented the meiotic arrest and cAMP accumulation mediated by follicle-stimulating hormone (FSH) in intact complexes. Therefore, inhibition of oocyte phosphodiesterase appears to be one mechanism by which the purines could maintain meiosis-arresting levels of cAMP.

1-Methyl-3-isobutylxanthine↗

Statistics in physiology and pharmacology: a slow and erratic learning curve.

1. Learning how to apply statistical analyses to the results of experimental or clinical studies may take a lifetime of trial (and sometimes error), as it has done in the author's case. There is no evidence that biomedical investigators of the present generation are on a steeper learning curve. Gross misunderstandings of the purpose and functions of statistical analysis are apparent in applications to research grant-giving bodies and ethics committees, in manuscripts submitted to journals and sometimes in published papers. 2. Although estimation of minimal group (sample) size for a given power is an essential step in planning clinical studies, it seems to be used rarely in laboratory experimental work. This is despite exhortations to restrict the number of animals used to a minimum. 3. Most investigators use hypothesis testing to analyse their results, but their understanding of the meaning of the resultant P-values is slight. 4. A flaw found almost universally in biomedical manuscripts is to make multiple inferences from the results of a single study. The goal of statistical analysis is to maintain the familywise type I error rate (risk of false-positive inference) at a predetermined level (usually 5%). But, when multiple inferences are made from the same experiment, the risk of false-positive error is inflated. There are two solutions to this problem: (i) use a multiple comparison procedure to control the familywise type I error rate; and (ii) test a single, global hypothesis. 5. Biomedical investigators have been quick to acquire computer statistics software and to use it to analyse their experiments. However, they have been slow to recognize the limitations of this software. These include: (i) inadequate documentation of routines, so that neither the user nor the reader of published papers can be sure how the tests have been executed; (ii) flawed algorithms for the execution of statistical procedures; and (iii) failure to recognize that the best software for their purposes is that which takes them just beyond their statistical horizons. 6. The obvious solution to these difficulties is to recruit a biomedical statistician into every research group, at a relatively trivial cost. However, properly qualified biostatisticians are in desperately short supply in Australia. It follows that research groups, national grant-giving agencies and academic institutions must make provision for the proper training and subsequent employment of biostatisticians.

Data Interpretation, Statistical↗