Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “hypothesis testing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

[Problems of clinical trials as viewed in the perspective of decision theory].

In this paper, the current clinical trial based on the Neyman-Pearson's principle of hypothesis testing is reviewed in the perspective of decision theory, its shortcomings are revealed and alternative methods proposed. In the current trial design, the alpha error is arbitrarily fixed at a level (e.g., 0.05), even lower than the beta error level (0.1 or 0.2). This results in failure with smaller trial size to identify the better treatment, dose-patient interaction, or optimal patient stratum for treatment. If on the other hand, the most harmless alpha error is ignored, the total number of patients expected to be cured can be substantially increased due to earlier identification of the better treatment. In case treatment, A has some disadvantages compared with B in terms of cost, toxicity etc., then the null hypothesis should be changed to "the cure rate of A is greater than that of B by delta", an adjustment for the disadvantages, which can be estimated by utility analysis. The loss to the society and target population may be further increased if improper regimen is provided including omission of the best treatment arm, suboptimal treatment intensity, selection of improper patient stratum, failure to individualize treatment plan (inflexible regimen). In order to maximize the benefit of our patients and society, radical changes of the current clinical trial design is needed using decision analytical approach.

Antineoplastic Agents↗

On the nature of insight solutions: evidence from skill differences in anagram solution.

According to the Gestalt psychologists, problem solutions that pop into mind suddenly with no awareness of the process by which they were generated are objectively as well as subjectively sudden. Thus, such pop-out solutions are qualitatively different from search solutions, which are constructed incrementally. The authors tested this claim in the domain of anagram solution. Experiment 1 documented that anagrams yield pop-out solutions, especially among highly skilled solvers. The results of Experiment 2 indicated that both pop-out and search solutions depended on the gradual accumulation of partial information, contrary to the Gestalt view of problem solving. Nevertheless, some aspects of the Experiment 2 results, as well as new analyses of an anagram study reported elsewhere, suggest that there may in fact be a qualitative difference between pop-out and search solutions. In particular, pop-out solutions may result from parallel processing of the constraints on the rearranged order of the anagram letters, whereas search solutions may result from a serial hypothesis-testing procedure. Like dynamite, the insightful solution explodes on the solver's cognitive landscape with breathtaking suddenness, but if one looks closely, a long fuse warns of the impending reorganization. (Durso, Rea, & Dayton, 1994, p. 98)

Aptitude↗

Methods for quality-of-life studies.

Methodologies involving the use of quality-of-life patient outcomes in observational and interventional studies of health are drawn from a large and diverse field of research methods. The multidimensional way in which quality of life is conceptualized will affect the way it is measured and the complexity of the measurement. At the earliest stages of research, one must rely on methods common to the fields of tests and measurement, survey research, psychometrics and sociometrics to measure constructs that are not directly observable. Indices measuring performance can either focus on the scale's ability to perform in noninterventional, cross-sectional studies or interventional, longitudinal studies. Indices of stability, internal consistency, responsiveness with respect to true changes in quality of life, and sensitivity to treatment effects can be used to assess the scale's adequacy as a dependent variable of interest. Respondent variability can occur due to factors such as different reporters (patient, spouse, physician), the manner and form of administration (long form vs short form; self-administration vs interview) and the assessment environment (clinic, home). Finally, since quality-of-life research often involves inferential statistics and hypothesis testing, the statistical and epidemiologic principles of good study design should be followed. In addition, one should account for the reliability, responsiveness, and the sensitivity of the scale when designing the scientific hypotheses, and should specifically address the meaning of quality-of-life effect sizes by interventional-based validation. Design considerations must address the statistical issues of power, the determination of effect sizes through validation by external criteria, longitudinal data, effects of withdrawal and early termination, ceiling and floor effects, and heterogeneity of responsiveness and sensitivity among individuals. The problem of estimating quality-of-life summary parameters for use in pharmacoeconomic models is receiving increasing attention in this era of health-care reform and fiscal restraint. While medical decision theory has used cost-effectiveness models and quality-adjusted life years since the early 1970s, estimation of population parameters to differentiate among different medical interventions is relatively new. The assessment of the patient outcomes associated with medical interventions in terms of the risks, benefits and costs will clearly be a major focus of health-care reform. Development of new methodologies in quality-of-life research should build upon the strong foundation already established in the areas of clinical research, epidemiology, biostatistics, economics and behavioral science.(ABSTRACT TRUNCATED AT 400 WORDS)

Bias↗

Bond strengths of new simplified dentin-enamel adhesives.

PURPOSE: To compare the in vitro shear bond strengths (SBS) of five simplified dentin adhesives. The tested hypothesis was that the recently introduced simplified adhesive systems would have similar or higher SBS than an existing simplified acetone-based adhesive used as a control. MATERIALS AND METHODS: 100 flat bonding sites were polished to 600-grit on the labial surface of bovine incisors mounted in acrylic resin. 50 teeth were ground to expose enamel, while the remaining 50 specimens were prepared to expose middle dentin. The specimens were randomly divided into five equal groups to be treated with simplified dentin adhesives: Dentastic Uno, EasyBond, Gluma One Bond, One Coat Bond, and One-Step (control). A composite post was bonded to each treatment area. After thermo-cycling, enamel and dentin shear bond strengths were determined using an Instron testing machine and the data were submitted to statistical analyses. RESULTS: Mean enamel bond strengths ranged from 14.6-28.4 MPa. One Coat Bond had the highest mean enamel SBS, but it was not significantly higher than those of Gluma One Bond and Dentastic Uno. EasyBond and One-Step had statistically similar mean enamel SBS and these were significantly lower than the mean enamel SBS of the other three adhesives. For dentin, mean SBS ranged from 14.8-21.7 MPa. Dentastic Uno had the highest mean dentin SBS, but it was not significantly greater than those of One Coat Bond and Gluma One Bond. Although One Step had the lowest mean dentin SBS, it was not significantly different from those of either EasyBond or Gluma One Bond.

Analysis of Variance↗

Reliability and validity of the self-efficacy and outcome expectations for osteoporosis medication adherence scales.

PURPOSE: The purpose of this study was to test the reliability and validity of the self-efficacy and outcome expectations for osteoporosis medication adherence measures (SEOMA and OEOMA). DESIGN: This was a descriptive study involving a single face-to-face interview. SAMPLE: The study included 152 older adults with a mean age of 85.7 (+) 5.5 years, the majority of whom were Caucasian (99%), female (74%), and unmarried (75%). METHODS: In addition to the SEOMA and OEOMA measures, demographic information (age, gender, and marital status) and other health behaviors (exercise and osteoporosis medication use) were explored. RESULTS: There was evidence of reliability of the SEOMA and OEOMA based on internal consistency and R values. Evidence of the validity of the SEOMA and OEOMA measures was based on confirmatory factor analysis and hypothesis testing. CONCLUSION: This study is an important first step to developing reliable and valid measures of self-efficacy and outcome expectations for adherence to osteoporosis medications. IMPLICATIONS FOR NURSING PRACTICE: The SEOMA and OEOMA can be used to evaluate self-efficacy and outcome expectancy beliefs related to osteoporosis medication use in older adults and interventions developed to strengthen those beliefs and improve medication adherence.

Aged↗

Testing the reliability and validity of the Self-Efficacy for Exercise scale.

BACKGROUND: The measure for self-efficacy barriers to exercise was developed for adults and revised on the basis of quantitative and qualitative research with older adults so it would be more appropriate for that age group. OBJECTIVES: To test the reliability and validity of the Self-Efficacy for Exercise (SEE) Scale. METHODS: Initial reliability and validity testing was performed using a sample of 187 older adults living in a continuing care retirement community. The average age of the participants was 85 +/- 6.2 years, and most were White (98%), female (82%), and unmarried (80%). Face-to-face interviews were completed and included the SEE, the 12-item Short Form Health Survey (SF-12), and the Expected Outcomes and Barriers for Habitual Exercise scale. Exercise activity was based on verbal report of participation in aerobic exercise (walking, swimming, biking, or jogging). RESULTS: There was sufficient evidence of internal consistency (alpha = 0.92), and a squared multiple correlation coefficient using structural equation modeling provided further evidence of reliability (R2 ranged from 0.38 to 0.76). There was evidence of validity of the measure based on hypothesis testing: Mental and physical health scores on the SF-12 predicted efficacy expectations, and efficacy expectations predicted exercise activity. Lambda X estimates (all estimates > or = 0.81) provided further evidence of validity. CONCLUSION: Preliminary testing provided evidence for the reliability and validity of the SEE scale. Future testing of the scale needs to be done with young old adults and subjects from different socioeconomic and cultural groups.

Aged↗

Effect of hydration variability on hybrid layer properties of a self-etching versus an acid-etching system.

The hypothesis tested in this study was that the self-etching system (Clearfil SE Bond, CSE) is less sensitive to surface moisture variability than the system that uses a separate acid-etching step (Single Bond, SB). Eighteen dentin specimens were bonded to composite using CSE or SB. Three different surface moisture conditions per bonding type (overwet, w; dry, d; and visibly moist, n [normal]) were applied prior to bonding dentin to composite. One cross section of each sample was analyzed with lines of nanoindentations crossing perpendicular to the bonding interface. An additional set of bonded samples was fixed and cross sectioned before the hybrid layer thickness was measured in scanning electron microscopy. The nanoindentations revealed significant differences in indentation modulus (E(i)) and hardness (H) for the hybrid layer comparing SBn, E(i) = 2.7(+/-1.6); H = 0.24(+/-0.1) GPa with SBd, E(i) = 0.9(+/-0.7); H = 0.9(+/-0.05) GPa, respectively, while CSE showed no differences among the groups. A significantly greater demineralized zone below the hybrid layer was found for SBd. The hybrid layer was wider for both CSEd and SBd. In conclusion the hypothesis was verified; CSE exhibited no significant changes of hybrid layer properties (E(i), H) at different hydration conditions, while SB had significant differences, especially after air-drying.

Acid Etching, Dental↗

The orderly progression of melanoma nodal metastases.

OBJECTIVE: The aim of this study was to determine the order of melanoma nodal metastases. SUMMARY BACKGROUND DATA: Most solid tumors are thought to demonstrate a random nodal metastatic pattern. The incidence of skip nodal metastases precluded the use of sampling procedures of first station nodal basins to achieve adequate pathological staging. Malignant melanoma may be different from other malignancies in that the cutaneous lymphatic flow is better defined and can be mapped accurately. The concept of an orderly progression of nodal metastases is radically different than what is thought to occur in the natural history of metastases from most other solid malignancies. METHODS: The investigators performed preoperative and intraoperative mapping of the cutaneous lymphatics from the primary melanoma in an attempt to identify the "sentinel" lymph node in the regional basin. All patients had primary melanomas with tumor thicknesses > 0.76 mm and were considered candidates for elective lymph node dissection. The sentinel lymph node was harvested and submitted separately to pathology, followed by a complete node dissection. The null hypothesis tested was whether nodal metastases from malignant melanoma occurred in equal proportions among sentinel and nonsentinel nodes. RESULTS: Forty-two patients met the criteria of the protocol based on prognostic factors of their primary melanoma. Thirty-four patients had histologically negative sentinel nodes, with the rest of the nodes in the basin also being negative. Thus, there were no skip metastases documented. Eight patients had positive sentinel nodes, with seven of the eight having the sentinel node as the only site of disease. In these seven patients, the frequency of sentinel nodal metastases was 92%, whereas none of the higher nodes had documented metastatic disease. Nodal involvement was compared between the sentinel and nonsentinel nodal groups, based on the binomial distribution. Under the null hypothesis of equality in distribution of nodal metastases, the probability that all seven unpaired observations would demonstrate that involvement of the sentinel node is 0.008. CONCLUSIONS: The data presented demonstrate that nodal metastases from cutaneous melanoma are not random events. The sentinel lymph nodes in the lymphatic basins can be mapped and identified individually, and they have been shown to contain the first evidence of melanoma metastases. This information can be used to revolutionize melanoma care so that only those patients with evidence of nodal metastatic disease are subjected to the morbidity and expense of a complete node dissection. Because sentinel node histology accurately reflects the histology of the remainder of the lymphatic basin, information gained from the sentinel node biopsy can be used as a prognostic factor for melanoma. These findings demonstrate effective pathologic staging, no decrease in standards of care, and a reduction of morbidity with a less aggressive, rational surgical approach.

Adolescent↗

Is cytochrome P450 2C9 genotype associated with NSAID gastric ulceration?

AIMS: The aim of this study was to explore whether genetic variation of cytochrome P450 2C9 (CYP2C9) contributes to NSAID-associated gastric ulceration. The hypothesis tested was that CYP2C9 poor metabolizer genotype would predict higher risk of gastric ulceration in patients on NSAIDs that are metabolized by CYP2C9, due to higher plasma NSAID concentrations. METHODS: Peripheral blood DNA samples from 23 people with a history of gastric ulceration attributed to NSAIDs metabolized by CYP2C9, and from 32 people on NSAIDs without gastropathy, were analysed to determine CYP2C9 genotype. RESULTS: The following genotypes were found: *1/*1 (wild type) in 70% of cases and 58% of controls, *1/*2 in 17% of cases and 29% of controls, *1/*3 in 13% of cases and 13% of controls. The difference between case and control nonwild-type genotype frequency was 11.5% (95% CI -14,37%), with the direction of the difference being against the hypothesis. No individuals with homozygote poor metaboliser genotype were identified. The differences in genotype frequencies between the two groups were not significant and the frequencies were similar to those in a large published population study. Ninety-five percent binomial confidence interval analysis confirms that there is no apparent clinically significant relationship between CYP2C9 genotype and risk of gastric ulceration although a small difference in risk in poor metabolizers cannot be excluded. CONCLUSIONS: These results do not support the hypothesis that gastric ulceration resulting from NSAID usage is linked to the poor metabolizing genotypes of CYP2C9.

Adult↗

The distribution of health care costs and their statistical analysis for economic evaluation.

OBJECTIVE: Where patient level data are available on health care costs, it is natural to use statistical analysis to describe the differences in cost between alternative treatments. Health care costs are, however, commonly considered to be skewed, which could present problems for standard statistical tests. This review examines how authors report the distributional form of health care cost data and how they have analysed their results. METHOD: A review of cost-effectiveness studies that collected patient-level data on health care costs. To supplement the review, five datasets on health care costs are examined. Consideration is given to the use of parametric methods on the transformed scale and to non-parametric methods of analysing skewed cost data. RESULTS: Since economic analysis requires estimation in monetary units, the usefulness of transformation-based methods is limited by the inability to retransform cost differences to the original scale. Non-parametric rank sum methods were also found to be of limited use for economic analysis, partly due to the focus on hypothesis testing rather than estimation. Overall, the non-parametric approach of bootstrapping was found to offer a useful test of the appropriateness of parametric assumptions and an alternative method of estimation where those assumptions were found not to hold. CONCLUSIONS: Guidelines for the analysis of skewed health care cost data are offered.

Cost-Benefit Analysis↗

A 2-D model of flow-induced alterations in the geometry, structure, and properties of carotid arteries.

Evidence from diverse investigations suggests that arterial growth and remodeling correlates well with changes in mechanical stresses from their homeostatic values. Ultimately, therefore, there is a need for a comprehensive theory that accounts for changes in the 3-D distribution of stress within the arterial wall, including residual stress, and its relation to the mechanisms of mechanotransduction. Here, however, we consider a simpler theory that allows competing hypotheses to be tested easily, that can provide guidance in the development of a 3-D theory, and that may be useful in modeling solid-fluid interactions and interpreting clinical data. Specifically, we present a 2-D constrained mixture model for the adaptation of a cylindrical artery in response to a sustained alteration in flow. Using a rule-of-mixtures model for the stress response and first order kinetics for the production and removal of the three primary load-bearing constituents within the wall, we illustrate capabilities of the model by comparing responses given complete versus negligible turnover of elastin. Findings suggest that biological constraints may result in suboptimal adaptations, consistent with reported observations. To build upon this finding, however, there is a need for significantly more data to guide the hypothesis testing as well as the formulation of specific constitutive relations within the model.

Animals↗

Movements of the pelvis and lumbar spine during walking in people with acute low back pain.

BACKGROUND AND PURPOSE: Little is known about how acute low back pain affects pelvic and lumbar movements during walking. The aim of the present study was to determine if measurement of the amplitude of the angular movements of the pelvis and lumbar spine during walking is useful in the evaluation of people with acute low back pain. METHOD: The study used a repeated-measures and correlational design; 11 individuals with low back pain (tested in the acute phase and six weeks later when symptoms had resolved) and matched control subjects were tested during treadmill walking. A video-analysis system was used to measure the amplitudes of movements of the pelvis and lumbar spine during walking. Pain level was measured with a visual analogue scale (VAS). RESULTS: During walking, movements of the pelvis (axial rotation) and lumbar spine (lateral flexion) were reduced in people with acute back pain compared with the resolved condition, but were not different from a group without a history of back pain. The amplitudes of the frontal plane movements of the pelvis and lumbar spine were negatively correlated with the intensity of pain (r(s) = -0.74). CONCLUSIONS: Measurement of pelvic and lumbar movements during walking is unlikely to have useful clinical applications for individuals, or when discriminating between impaired and unimpaired people, but may be applied to groups for hypothesis testing in evaluating change in back pain over time. An hypothesized model to explain the observed movements has been proposed.

Acute Disease↗

Development of a standardized animal model for the study of alkali ingestion.

The purpose of this study was to develop an animal model that would grade in-vivo therapeutic modality testing for caustic ingestion. Caustic substances are found in many household items (eg detergents, bleaches, pipe cleaners) and pose a serious threat to health if ingested accidentally or intentionally with resulting injuries including immediate death or chronic debilitating morbidity. This study used 5, 3.8 or 1.8% sodium hydroxide (NaOH) to determine macro/microscopic injury at 10, 30, or 60 minutes. Macroscopic grading was based on gross evaluation of denudation of mucosa, edema, hyperemia, hemorrhage, ulcerations and necrosis. Microscopic grading was based on epithelial viability, cornified epithelial cell differentiation, granular cell differentiation, epithelial cell nuclei, muscle cell viability and muscle cell nuclei. Product concentration was shown a more significant predictor of injury than time of exposure. The grading system presented should provide a reliable method of producing and grading alkaline ingestions for future treatment hypothesis testing.

Administration, Oral↗

Regulation of hepatocyte thyroxine 5'-deiodinase by T3 and nuclear receptor coactivators as a model of the sick euthyroid syndrome.

The syndrome of nonthyroidal illness, also known as the sick euthyroid syndrome, is characterized by a low plasma T3 and an "inappropriately normal" plasma thyrotropin in the absence of intrinsic disease of the hypothalamic-pituitary-thyroid axis. The syndrome is due in part to decreased activity of type I iodothyronine 5'-deiodinase (5' D-I), the hepatic enzyme that converts thyroxine to T3 and that is induced at the transcriptional level by T3. The hypothesis tested is that cytokines decrease T3 induction of 5' D-I, resulting in decreased T3 production and hence a further decrease in 5' D-I. The proposed mechanism is competition for limiting amounts of nuclear receptor coactivators between the 5' D-I promoter and the promoters of cytokine-induced genes. Using primary cultures of rat hepatocytes, we demonstrate that interleukins 1 and 6 inhibit the T3 induction of 5' D-I RNA and enzyme activity. This effect is at the level of transcription and can be partially overcome by exogenous steroid receptor coactivator-1 (SRC-1). The physical mass of endogenous SRC-1 is not affected by cytokine exposure, and exogenous SRC-1 does not affect 5' D-I in the absence of cytokines. The data support the hypothesis that cytokine-induced competition for limiting amounts of coactivators decreases hepatic 5' D-I expression, contributing to the etiology of the sick euthyroid syndrome.

Animals↗

Repeated measurement designs with random selections.

We consider repeated measurement designs with a single group as well as with multiple groups. Conventionally, the repeated measurements are made at fixed, often evenly spaced times. The first and second moment conditions needed for an exact F test are not satisfied in general, but with random permutation of the times, a probability measure results that does satisfy these moment conditions, unconditionally. As a consequence, unfortunately, we lose the asymptotic normality that is also essential to justify an F test. We introduce a class of designs and estimators for which both the moment conditions and asymptotic normality are satisfied. The times are the same for all subjects and they are chosen in such a way that they are mutually independent and identically distributed. It is important to understand that our conditions are satisfied only unconditionally--that is, if we do not condition by the randomly chosen times. There is a small probability that the randomly sampled times could be very unevenly allocated. Since we are not conditioning by the times, strictly speaking, this does not matter. It is possible to impose a mild condition of "spread," as we show, at small cost in terms of the P-value. Our method of design makes it possible to do regression analysis, interval estimation, and hypothesis testing on the mean response as a function of time. Because we have mutual independence and identical distributions of times we can form confidence bands.

Animals↗

Various criteria in the evaluation of biomedical named entity recognition.

BACKGROUND: Text mining in the biomedical domain is receiving increasing attention. A key component of this process is named entity recognition (NER). Generally speaking, two annotated corpora, GENIA and GENETAG, are most frequently used for training and testing biomedical named entity recognition (Bio-NER) systems. JNLPBA and BioCreAtIvE are two major Bio-NER tasks using these corpora. Both tasks take different approaches to corpus annotation and use different matching criteria to evaluate system performance. This paper details these differences and describes alternative criteria. We then examine the impact of different criteria and annotation schemes on system performance by retesting systems participated in the above two tasks. RESULTS: To analyze the difference between JNLPBA's and BioCreAtIvE's evaluation, we conduct Experiment 1 to evaluate the top four JNLPBA systems using BioCreAtIvE's classification scheme. We then compare them with the top four BioCreAtIvE systems. Among them, three systems participated in both tasks, and each has an F-score lower on JNLPBA than on BioCreAtIvE. In Experiment 2, we apply hypothesis testing and correlation coefficient to find alternatives to BioCreAtIvE's evaluation scheme. It shows that right-match and left-match criteria have no significant difference with BioCreAtIvE. In Experiment 3, we propose a customized relaxed-match criterion that uses right match and merges JNLPBA's five NE classes into two, which achieves an F-score of 81.5%. In Experiment 4, we evaluate a range of five matching criteria from loose to strict on the top JNLPBA system and examine the percentage of false negatives. Our experiment gives the relative change in precision, recall and F-score as matching criteria are relaxed. CONCLUSION: In many applications, biomedical NEs could have several acceptable tags, which might just differ in their left or right boundaries. However, most corpora annotate only one of them. In our experiment, we found that right match and left match can be appropriate alternatives to JNLPBA and BioCreAtIvE's matching criteria. In addition, our relaxed-match criterion demonstrates that users can define their own relaxed criteria that correspond more realistically to their application requirements.

Algorithms↗

Conspecific urine marking in male-female pairs of laboratory rats.

Male-female pairs of rats were observed during social interactions for conspecific markings where an animal deposits urine on the body of a second animal. Injections of sodium fluorescein were used to change urinary color and provide a visible mark on the back of a conspecific. The general hypothesis tested in the three experiments was that a female would mark males differently dependent upon her hormonal status and upon the relative hormonal integrity of the males. Results of Experiment 1 were that males marked females more frequently than vice versa and female marking decreased during her estrus. Moreover, a diestrous female marked an aggressive male more than she marked a non-aggressive male (Experiment 2), and a diestrous female marked a castrated male receiving 800 micrograms testosterone propionate (TP) injections more copiously than she marked a castrated male receiving 200 micrograms TP injections (Experiment 3). Finally, aggressive and non-aggressive males marked females similarly, though the 800 micrograms TP males marked females more than the 200 micrograms TP males. These data were compared with findings from research on environmental marking in rodents, and an hypothesis was suggested that female rats use conspecific marking to identify and select males with preferred behavioral and endocrine characteristics.

Aggression↗

Beliefs about rape and women's social roles.

The hypothesis tested was that beliefs about rape that place women at a disadvantage are positively related to beliefs that restrict the rights and roles of women in our society. Two scales, the R scale and the W scale, based on a survey of beliefs about rape (Feild, 1978) and the Attitudes Toward Women Scale (Spence and Helmreich, 1972), were administered as a single instrument. Subjects included 432 female undergraduates, 140 male undergraduates, 114 employed women, and 76 employed men. The latter two groups were predominantly from managerial, technical, and professional occupations. Product moment correlations between responses on the R scale and responses on the W scale were calculated for total scores as well as for three factors: women's responsibility and causal role in rape; role of consent in rape; and rapist's role and motivation. Correlations consistently supported the hypothesis for all four groups.

Attitude↗