Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

The diagnostic validity of the Athens Insomnia Scale.

OBJECTIVE: To provide documentation for the diagnostic validity of the Athens Insomnia Scale (AIS), a self-assessment psychometric tool which has previously shown high consistency, reliability and external validity for the evaluation of the intensity of sleep difficulty. METHODS: The AIS was administered to a total of 299 subjects (105 primary insomniacs, 100 psychiatric outpatients, 44 psychiatric inpatients and 50 nonpatient controls) who were also assessed for the ICD-10 diagnosis of "nonorganic insomnia" blindly in terms of the AIS scores. RESULTS: 176 subjects were identified as insomniacs and 123 as noninsomniacs. Logistic regression of AIS total score against the ICD-10 diagnosis of insomnia demonstrated that a score of 6 is the optimum cutoff based on the balance between sensitivity and specificity. When diagnosing individuals with a score of 6 or higher as insomniacs, the scale presents with 93% sensitivity and 85% specificity (90% overall correct case identification). For this cutoff score, in the general population, the scale has a positive predictive value (PPV) of 41% and a negative predictive value (NPV) of 99%. For the same cutoff score, among unselected psychiatric patients, the PPV was found to be 86% and the NPV 92%. Other cutoff scores can be also considered, however, depending on the importance of avoiding false positive or false negative results; for example, for a cutoff score of 10, the PPV in the general population reaches about 90% without the NPV becoming lower than 94%. CONCLUSION: The AIS can be utilized in clinical practice and research, not only as an instrument to measure the intensity of sleep-related problems, but also as a screening tool in reliably establishing the diagnosis of insomnia.

Adult↗

Cognitive failures and performance differences: validation studies of a German version of the cognitive failures questionnaire.

A German version of Broadbent's Cognitive Failures Questionnaire was developed because there was no comparable German instrument. As external validity criteria observable failures in everyday situations were used. The positive correlations were significant but moderate in size. For construct validation, inventories of trait anxiety, coping styles and action control were used. Most of the correlations were significant and in the expected direction. A maximum likelihood factor analysis of the questionnaire suggested that the utility of the total score may be doubtful. The unidimensionality of the construct requires further investigation.

Accident Proneness↗

A structured diagnostic interview for hypochondriasis. A proposed criterion standard.

We developed a structured diagnostic interview for DSM-III-R hypochondriasis (SDIH) that is the first such clinician-administered instrument. The SDIH was administered to 88 general medical outpatients who scored above a predetermined cutoff on a hypochondriacal symptom questionnaire, and to 100 comparison patients randomly chosen from among those below the cutoff. Using the joint assessment method, interrater agreement on the DSM-III-R diagnostic criteria was 88% to 97% and agreement on the diagnosis was 96%. Concurrent validity was suggested by a significant correlation between the interview and the primary care physicians' ratings of hypochondriasis. A measure of external validity was demonstrated in that several clinical characteristics thought to be ancillary features of hypochondriasis were significantly more prevalent in interview-positive patients than in interview-negative patients. Finally, the SDIH appeared to have discriminant validity in that patients diagnosed as hypochondriacal had several other clinical features that distinguished them from the patients who scored above the cutoff on hypochondriacal symptomatology, but failed to be diagnosed as hypochondriacal with the SDIH.

Ambulatory Care↗

Clinical trial design in schizophrenia: implications for clinical decisions.

PURPOSE OF REVIEW: Results from clinical trials do not necessarily provide information for decisions in clinical practice. This review aims to present strengths and limitations of different methodological types of clinical trials and to offer an overview of how knowledge from clinical trails can be distilled for clinical practice. Selected key questions in the treatment of schizophrenia are presented, with a focus on the possibilities and restrictions of translating trial results into real-world practice. RECENT FINDINGS: Randomized controlled trials are the gold standard for proving efficacy of a diagnostic or therapeutic procedure. They have a high degree of internal validity and a clear-cut message when conducted to good-quality standards but suffer from a lack of generalizability (external validity). Effectiveness studies evaluate effects of treatments under conditions approximating usual care. They may include patient-centred outcomes or health economic evaluations. According to the type of trial, specific problems arise in the interpretation of results. Typical examples are given for the treatment of acute exacerbations of schizophrenia, for relapse prevention and for the treatment of cognitive impairment. SUMMARY: Clinical decisions have to be made upon the best knowledge. Therefore, well conducted studies addressing all major issues from all relevant perspectives are needed. The assessment of a treatment regimen for clinical utility requires both efficacy and effectiveness studies. An understanding of the design, analysis and conventions of both study types is essential for the interpretation of results and their translation to the clinical decision-making process.

Journal Article↗

The proposition: an insight into research.

Propositions form the basis for scientific research. The validity of a research study is, to a large extent, evaluated on the criteria of its propositions. For internal validity, study propositions provide information regarding precision of definitions, measurements, associations, confounding factors etc. that are considered in research. While for external validity, propositions form the premise for the deduction of inferences. The aim of this article is to help readers understand the propositions that are made in research. This article discusses those propositions, which are relevant to medical research.

Causality↗

Establishing the diagnostic validity of premenstrual dysphoric disorder using rasch analysis.

Premenstrual Dysphoric Disorder (PMDD) has remained in appendices of the last two editions of The Diagnostic and Statistical Manual of Mental Disorders due to lack of empirical study. Items included in its set of research criteria are considered tentative pending evidence of diagnostic validity. The present study attempts to establish the construct validity of the PMDD criteria using the Rasch method to analyze the validity of individual items as contributors to the diagnosis, in contrast to the usual but less precise approach of using an external validator to establish the diagnostic utility of psychiatric conditions. Analysis of which items best differentiate participants with and without PMDD provides an idea of the relative ability of these items to distinguish PMDD. It is recommended that the areas of anger/irritability, depressed mood, and problems in interpersonal functioning be expanded in further studies and corresponding items added to symptom checklists.

Adolescent↗

Alcohol surveys with high and low coverage rate: a comparative analysis of survey strategies in the alcohol field.

Two Swedish alcohol surveys were compared in a search for a reasonable explanation of the large difference in their coverage rates, namely 75% and 28%. In many respects both surveys conducted in the late 1980s by large, well-known institutes, are of a similar type with rather large samples of Swedes. The technique used in the survey with a very high coverage rate (Survey A) takes into consideration the actual drinking pattern of the population studied (i.e., the concentration of drinking on weekends). By dividing a "normal week's consumption" into four units (Monday-Thursday, Friday, Saturday and Sunday), the technique allows one to average periods with varying drinking habits. In the survey with a low coverage rate (Survey B) a "normal week's consumption" was not so divided. A test of internal validity within Survey A underlined the general finding that its higher coverage rate was due to this division. A test of the external validity at aggregate level did not support assumptions about "telescoping" effects in A. Both A and B had a normal week as a basis of measurement for investigating typical drinking habits. The literature concerning differences in coverage rates focuses on the measurement of modal habits versus mean habits. The main explanation of differences is that methods that focus on modal habits (i.e., the Quantity-Frequency Scale) generate a lower coverage rate than do methods that elicit the arithmetic mean (i.e., the last-week recall). Since A and B both belong to the former type of scale, this does not explain our results.

Adolescent↗

[Quality of life and multiple sclerosis: validation of the french version of the self-questionnaire (SEP-59)].

We conducted a prospective study among 166 multiple sclerosis (MS) patients (103 from an university hospital, 63 from a MS rehabilitation center) to assess the properties of the French version of the Multiple Sclerosis Quality Of Life - 54 items (MS QOL-54) which combines the MOS SF36 together with MS specific items. The SF-36 had been translated into French through the IQOLA project. We translated and adapted the MS specific items with the help of three different teams. The translation into French has an addition of five items, because we kept the MS specific items of an earlier unpublished form. Acceptability is excellent with a response rate over 90p.100. Test-retest reliability is good except for the "role limitation-emotional" scale of the SF-36. Construct validity, based on factor analysis, shows no change in the SF-36 internal consistency and the specific items provided their own information. External validity, tested against both medical (Expanded Disability Status Scale, Kurtzke scale, Mini-Mental-State and disease stage) and rehabilitation (Functional Independence Measure) parameters is excellent. The French MS QOL questionnaire contains 59 items including both the SF-36 and the MS QOL-54 items. This will permit international comparisons of MS patients' care and therapy.

Adolescent↗

[Clinical cardiology and physiology].

Recent concepts of cardiovascular pharmacotherapy are now mainly based on results of multicentre prospective studies performed in several countries worldwide. The main disadvantage of such studies is the heterogeneity of patient-groups compared, caused by ethnic differences, variability of criteria and diverse concomitant medication. To assert internal validity of data compared, substantial section of patients (80-99%) is often excluded before evaluation. Thus, there is always a lack of external validity--the data cannot be generalised and have only a limited predictive value for management of other groups or individual patients. Consequently, we still need reliable preclinical data from animal and human studies under well defined and uniform conditions. The importance of experimental and clinical physiology is demonstrated by well-established hemodynamic rules and pharmacodynamic correlations. Clinical physiology remains an indispensable foundation of clinical medicine which cannot be replaced by formal statistics.

Animals↗

Validation of predictors of intraprocedural stent thrombosis in the drug-eluting stent era.

Although predictors of acute intraprocedural stent thrombosis (IPST) in the drug-eluting stent era have been proposed, external validation is lacking. We thus analyzed the occurrence of IPST in the RECIPE study and found that, among 1,320 patients who underwent drug-eluting stent implantation, IPST occurred in 6 (0.5%), with in-hospital major adverse events in 4 (67%). IPST was predicted by number and total length of implanted stents, baseline minimal lumen diameter, and, in a pooled analysis that incorporated values from the present study and a previous study, use of elective glycoprotein IIb/IIIa inhibitors. Such results may provide useful information to guide prevention of this complication.

Acute Disease↗

[Impact of the availability of an external nuclear medicine service in the application of sentinel lymph node biopsy in breast cancer surgery].

INTRODUCTION: To perform sentinel lymph node biopsy (SLNB), nuclear medicine services that have previously undergone a validation phase are required. The aim of the present study was to analyze the possibility of performing this technique with a previously validated, external nuclear medicine service and to study its impact on the indication for radical axillary lymphadenectomy (RAL) and on length of postoperative hospital stay. PATIENTS AND METHODS: We performed a prospective study in a cohort of patients with breast cancer starting from the introduction of SLNB in our center, which was made possible by collaboration with an external nuclear medicine service that performed lymphoscintigraphy and sentinel node detection. Intraoperative detection was performed through a portable probe. The feasibility of the project and its clinical impact were analyzed, taking a reduction in the number of lymphadenectomies and length of hospital stay as endpoints. RESULTS: A total of 196 patients with 201 breast carcinomas were treated. The most frequent interventions were tumorectomy (TC) with SLNB in 124 patients (62%), and TC with SLNB and RAL in 62 patients (31%). Sentinel node visualization on lymphoscintigraphy was achieved in 187/201 carcinomas (93.1%) and sentinel nodes were detected during the intervention in 182/187 carcinomas (97.4%). Sentinel node detection in the internal mammary chain was achieved in 23/201 carcinomas (11.4%). RAL was avoided in 131 of the 201 carcinomas (65%). Days of postoperative hospital stay with or without RAL showed a mean difference of 1.8 days (3.1 vs. 1.3; P < .001). CONCLUSION: SLNB is feasible with the collaboration of an external nuclear medicine service. This technique avoids 65% of RAL and reduces length of postoperative stay by 1.8 days.

Adult↗

Common sense and figures: the rhetoric of validity in medicine (Bradford Hill Memorial Lecture 1999).

Austin Bradford Hill was once a friend to The Lancet, but, as occasionally happens, friends fall out. The great legacy of his association with the journal, however, was Principles of Medical Statistics. As each edition was succeeded by another--the first in 1937, the last in 1991--he seemed to shift his view about the influence of statistical method on clinical practice from one of assured certainty to one of modest advantage. That change paralleled a move away from an emphasis on the importance of internal validity in the randomized trial to one of understanding the inescapably practical significance of generalizability. Writers on medical research have explored notions of external validity in various ways. One view, for example, is to seek a close correlation between the participants in a clinical trial and patients seen in practice. The argument goes that such a correspondence has to be made before any decision can be taken about whether to apply the result of that trial to the clinical setting. Another view, first worked out by the American logician Charles Sanders Peirce, is that one must simply rely on the informed guess, based on a reasonable estimate of the limits of extrapolation. The tensions between and implications of these two different approaches are worked through using the example of coronary stents. A solution is, perhaps, to write explicit rules of interpretation that provide a framework for judging the strength of a claim to applicability. Five questions are posed, which try to lay a foundation for such a framework.

Clinical Medicine↗

Measurement of impulsivity: construct coherence, longitudinal stability, and relationship with externalizing problems in middle childhood and adolescence.

This study focused on the assessment of impulsivity in nonreferred school-aged children. Children had been participants since infancy in the Bloomington Longitudinal Study. Individual differences in impulsivity were assessed in the laboratory when children were 6 (44 boys, 36 girls) and 8 (50 boys, 39 girls) years of age. Impulsivity constructs derived from these assessments were related to parent and teacher ratings of externalizing problems across the school-age period (ages 7-10) and to parent and self-ratings of these outcomes across adolescence (ages 14-17). Consistent with prior research, individual measures of impulsivity factor-analyzed into subdimensions reflecting children's executive control capabilities, delay of gratification, and ability or willingness to sustain attention and compliance during work tasks. Children's performance on the main interactive task index, inhibitory control, showed a signficant level of stability between ages 6 and 8. During the school-age years, children who performed impulsively on the laboratory measures were perceived by mothers and by teachers as more impulsive, inattentive, and overactive than others, affirming the external validity of the impulsivity constructs. Finally, impulsive behavior in the laboratory at ages 6 and 8 predicted maternal and self-ratings of externalizing problem behavior across adolescence, supporting the long-term predictive value of the laboratory-derived impulsivity measures.

Adolescent↗

Obesity prevention: a proposed framework for translating evidence into action.

Obesity as a major public health and economic problem has risen to the top of policy and programme agendas in many countries, with prevention of childhood obesity providing a particularly compelling mandate for action. There is widespread agreement that action is needed urgently, that it should be comprehensive and sustained, and that it should be evidence-based. While policy and programme funding decisions are inevitably subject to a variety of historical, social, and political influences, a framework for defining their evidence base is needed. This paper describes the development of an evidence-based, decision-making framework that is particularly relevant to obesity prevention. Building upon existing work within the fields of public health and health promotion, the Prevention Group of the International Obesity Task Force (IOTF) developed a set of key issues and evidence requirements for obesity prevention. These were presented and discussed at an IOTF workshop in April 2004 and were then further developed into a practical framework. The framework is defined by five key policy and programme issues that form the basis of the framework. These are: (i) building a case for action on obesity; (ii) identifying contributing factors and points of intervention; (iii) defining the opportunities for action; (iv)evaluating potential interventions; and (v) selecting a portfolio of specific policies, programmes, and actions. Each issue has a different set of evidence requirements and analytical outputs to support policy and programme decision-making. Issue 4 was identified as currently the most problematic because of the relative lack of efficacy and effectiveness studies. Compared with clinical decision-making where the evidence base is dominated by randomized controlled trials with high internal validity, the evidence base for obesity prevention needs many different types of evidence and often needs the informed opinions of stakeholders to ensure external validity and contextual relevance.

Decision Making↗

A systematic review of the performance of methods for identifying carious lesions.

This systematic review evaluates evidence describing histologically validated performance of methods for identifying carious lesions. A search identified 1,407 articles, of which 39 were included that described 126 assessment of visual, visual/tactile, radiographic (film and digital), fiber optic transillumination, electrical conductance, and laser fluorescence methods. A subsequent update added four studies contributing 10 assessments. The strength of the evidence was judged to be poor for all applications, signifying that the available information is insufficient to support generalizable estimates of the sensitivity and specificity of any given application of a diagnostic method. The literature is problematic with respect to complete reporting of methods, variations in histological validation methods, the small number of in vivo studies, selection of teeth, small numbers of examiners, and other factors threatening both internal and external validity. Future research must address these problems as well as expand the range of assessments to include primary teeth and root surfaces.

Dental Caries↗

[Value of the interactional anger model: studies in management and competitive sports].

This study focuses on testing the interactional approach underlying the anger models of Novaco (1978) and Spielberger (1988). An additional goal was to demonstrate different types of situations that give rise to anger and different dimensions of anger response in two samples of 97 managers and 74 athletes. Results showed that an anger response could be differentiated into a physiological, a cognitive, and two behavioral components. For the latter, it was confirmed that the expression of anger could be categorized into two independent components, anger in and anger out. Classifying anger situations revealed only a limited situation specificity of the tendency toward anger. The results on the validity of the interactional approach depended on the methods of analysis chosen. If the percentage of variance components were interpreted in the direction of an interactional approach, the external validity coefficients in at least one sample could not be interpreted unequivocally in this direction.

Adult↗

[Interpretation of results of clinical trials in benign prostatic hyperplasia].

The random clinical trial (RCT) is the most suitable study to evaluate the treatment effectiveness in the benign prostatic hyperplasia (BPH). Although most of the urologists will not collaborate in a RCT development, they will treat BPH patients, so it is very important to know if a CRT in BPH is well designed and their conclusions are correct. The aim of this article is to give the basic elements of analysis that urologists need in order to evaluate the quality and the level of evidence of a RCT in BPH. This article emphasizes the three main elements of a RCT: to check if the study has been correctly performed (internal validity), to evaluate if the treatment achieves an important clinical improvement (relevance of the results) and the applicability of the results in our patients (external validity). The article shows that to analyse these elements common sense and clinical judgment are needed rather than statistical knowledge.

Data Interpretation, Statistical↗

Devil's Claw (Harpagophytum procumbens) as a treatment for osteoarthritis: a review of efficacy and safety.

BACKGROUND: Osteoarthritis (OA) is a highly prevalent musculoskeletal disorder. Conventional treatment (i.e., the use of nonsteroidal anti-inflammatory drugs-NSAIDs) is associated with well-documented adverse effects. Devil's Claw (Harpagophytum procumbens) a traditional South African herbal remedy used for rheumatic conditions, may be a safer treatment option. To date, 14 clinical trials have assessed its efficacy/ effectiveness in OA. AIM: To address the two main questions of importance to clinicians: (1) Does Devil's Claw work for the treatment of OA, and (2) Is it safe? METHODS: A review of the literature on Devil's Claw and OA from 1966 to 2006 was performed using multiple search databases, monographs, and citation tracking. Relevant trials in all languages were identified and included. Both internal validity (i.e., adequacy of the dosage and period of treatment for this condition, reporting of randomization, rates of dropout, blinding, and statistical analysis) and external validity (i.e., inclusion/ exclusion criteria, baseline characteristics of the study populations, trial setting, and the appropriateness of the outcome measures of the trials) were assessed. RESULTS: Fourteen studies were identified: eight observational studies; 2 comparator trials (1 open, the other randomized to assess clinical effectiveness); and 4 double-blinded, placebo-controlled, randomized controlled trials to assess efficacy. Many of the published trials lacked certain important methodological quality criteria. However, the data from the higher quality studies suggest that Devil's Claw appeared effective in the reduction of the main clinical symptom of pain. The assessment of safety is limited by the small populations generally evaluated in the clinical studies. From the current data, Devil's Claw appears to be associated with minor risk (relative to NSAIDs), but further long-term assessment is required. CONCLUSIONS: The methodological quality of the existing clinical trials is generally poor, and although they provide some support, there are a considerable number of methodologic caveats that make further clinical investigations warranted. The clinical evidence to date cannot provide a definitive answer to the two questions posed: (1) Does it work? And (2) is it safe? A definitive high-quality trial that addresses the necessary methodologic improvements noted is needed to answer these important clinical questions.

Analgesics↗