Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Some methodologic lessons learned from cancer screening research.

Credible and useful methodologic evaluations are essential for increasing the uptake of effective cancer screening tests. In the current article, the authors discuss selected issues that are related to conducting behavior change interventions in cancer screening research and that may assist researchers in better designing future evaluations to increase the credibility and usefulness of such interventions. Selection and measurement of the primary outcome variable (i.e., cancer screening behavior) are discussed in detail. The report also addresses other aspects of study design and execution, including alternatives to the randomized controlled trial, indicators of study quality, and external validity. The authors conclude that the uptake of screening should be the main outcome when evaluating cancer screening strategies; that researchers should agree on definitions and measures of cancer screening behaviors and assess the reliability and validity of these definitions and measures in different populations and settings; and that the development of methods for increasing the external validity of randomized designs and reducing bias in nonrandomized studies is needed.

Biomedical Research↗

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans↗

Popperian epidemiology and the logic of bi-conditional modus tollens arguments for refutational analysis of randomised controlled trials.

Popperian epidemiology is a biomedical science tool based on the hypothesis-deductive method and the falsifiability of scientific hypotheses. This article explores the applicability of the refutationist logic tools in the analysis of a randomised controlled trial (RCT), the randomised Aldactone evaluation study (RALES). This was carried out by using bi-conditional modus-tollens arguments of the type (i) P-then-Q(n) and (ii) Q(n)-If-X(P), X(P) being a set of potential falsifiers of Q(n) as part of the explicit falsity-content of P. In this model, P is the main hypothesis and Q(n) one or more logical predictions to be tested. The X(P) argument represents inclusion criteria, exclusion criteria and conditional criteria of the RCT so every P-then-X(P) argument should be fulfilled in canonical form to corroborate P-then-Q(n). Thus, falsifiability of a RCT would be determined by the empirical content of the conditional argument Q(n)-If-X(P) and its external validity would be determined by the empirical content of X(P). In this way it would be possible to mathematically assess the external validity of a RCT if the observational predicates of the X(P) argument in a given population are known. According to this popperian model, applicability of the RCT results to clinical practice implies transferring of all its empirical content, in other words, the totality of its truth and falsity contents. Thus, to ignore the explicit falsity-content of a RCT such as RALES may jeopardise its potential benefits in clinical practice as suggested by recent studies.

Clinical Trials as Topic↗

Using the Internet to conduct surveys of health professionals: a valid alternative?

OBJECTIVE: The purpose of this study was to examine whether Internet-based surveys of health professionals can provide a valid alternative to traditional survey methods. METHODS: (i) Systematic review of published Internet-based surveys of health professionals focusing on criteria of external validity, specifically sample representativeness and response bias. (ii) Internet-based survey of GPs, exploring attitudes about using an Internet-based decision support system for the management of familial cancer. RESULTS: The systematic review identified 17 Internet-based surveys of health professionals. Whilst most studies sampled from professional e-directories, some studies drew on unknown denominator populations by placing survey questionnaires on open web sites or electronic discussion groups. Twelve studies reported response rates, which ranged from nine to 94%. Sending follow-up reminders resulted in a substantial increase in response rates. In our own survey of GPs, a total of 268 GPs participated (adjusted response rate = 52.4%) after five e-mail reminders. A further 72 GPs responded to a brief telephone survey of non-respondents. Respondents to the Internet survey were more likely to be male and had significantly greater intentions to use Internet-based decision support than non-respondents. CONCLUSIONS: Internet-based surveys provide an attractive alternative to postal and telephone surveys of health professionals, but they raise important technical and methodological issues which should be carefully considered before widespread implementation. The major obstacle is external validity, and specifically how to obtain a representative sample and adequate response rate. Controlled access to a national list of NHSnet e-mail addresses of health professionals could provide a solution.

Attitude of Health Personnel↗

Screening for eating disorders and high-risk behavior: caution.

OBJECTIVE: The current study reviews the state of eating disorder screens. METHODS: Screens were classified by their purported screening function: identification of cases with (a) anorexia nervosa only; (b) bulimia nervosa only; (c) eating disorders in general; (d) partial syndrome, eating disorder not otherwise specified (EDNOS), or subclinical; (e) not a-d but at high risk. Information is presented on development, psychometric properties, and external validation (e.g., sensitivity, specificity, positive predictive values, and negative predictive values). RESULTS: Screens differ widely with regard to objective, psychometric properties and the validation methodology used. Most screens that identify cases are not appropriate for the identification of at-risk behaviors. Little data on the external validity of screens are available. DISCUSSION: Screens should be used with caution. A sequential procedure, in which subjects identified as being at risk during the first stage is followed by more specific diagnostic tests during the second stage, might overcome some of the limitations of the one-stage screening approach.

Feeding and Eating Disorders↗

Using prognostic models in clinical infertility.

The chance that a couple who have tried to conceive for 12 months will succeed without assisted conception treatment is still higher than the chance that the same couple will benefit from treatment. In this context, it is important to assess the chance that a treatment-independent or 'spontaneous' pregnancy will occur in a couple whose wish for a child is unfulfilled. Prognostic models can be useful in this assessment. In recent years, prognostic models have been published both for the occurrence of 'spontaneous' pregnancy and for pregnancy after in vitro fertilization. This article discusses the theoretical aspects of prognostic modelling and assesses whether the current prognostic models are good enough to justify their use in clinical practice. The performance of existing models for the prediction of spontaneous conception was found to be acceptable on internal as well as on external validation. However, the performance of the existing models predicting IVF outcome was found to be disappointing on the few occasions on which such external validation has been performed.

Journal Article↗

Predictive validity of the strain index in manufacturing facilities.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate its predictive validity in two manufacturing plants. While blinded to health outcomes, investigators analyzed the right and left sides of 28 single-task jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder occurred on the right or left side during the prior three years, that side was classified as "positive." If no such disorder was reported during the prior three years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant with an empirical odds ratio of 73.2. The sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.84, 0.47, and 1.00. Similar results were noted for the jobs--the empirical odds ratio was 106.6, and the sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.91, 0.75, and 1.00. While these results provide additional evidence of the Strain Index's external validity and predictive validity, it should be noted that these jobs involved the performance of single tasks.

Arm Injuries↗

Predictive validity of the Strain Index in turkey processing.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate the predictive validity of the Strain Index in one turkey processing plant. While blinded to health outcomes, investigators analyzed the right and left sides of workers in 28 jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder had occurred on the right or left side during the previous 3 years, that side was classified as "positive." If no such disorder was reported during the previous 3 years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant, with an odds ratio of 22.0. The sensitivity, specificity, positive predictive value, and negative predictive value were 0.86, 0.79, 0.92, and 0.65, respectively. Similar results were noted for the jobs--the odds ratio was 50.0, and the sensitivity, specificity, positive predictive value, and negative predictive value were 0.91, 0.83, 0.95, and 0.71. These results provide additional evidence of the external validity and predictive validity of the Strain Index.

Animals↗

Development and validation of a risk score for somatic erectile dysfunction: combined results from three cross-sectional surveys.

OBJECTIVES: Some men with erectile dysfunction (ED) have difficulties discussing their condition with their physicians. Existing screening and diagnostic tools for ED often require the administration of personal questions regarding the condition. We present a simple risk score to estimate the individual likelihood of somatic ED, based on age and existing health conditions. METHODS: Data from the Cologne Male Survey (n = 4396) were used to develop a multivariable logistic regression model for the individual ED likelihood. The regression equation was both internally and externally validated using data from a national study (Berlin study) and a multinational cross-sectional study (MALES study). RESULTS: A final regression equation including age, pelvic surgery, diabetes mellitus, arterial circulatory disorder, heart disease, smoking, and hypertension reached an area under the receiver operating characteristic curve of 0.84 (0.5 means random and 1.0 perfect discrimination). Internal validation did not indicate any relevant overfit and the external validation results (national data: AUC = 0.75; multinational data: AUC = 0.67) are similar to those of other popular risk scores. CONCLUSIONS: The validated ED risk score developed from the regression equation can be used as a screening tool to identify patients who are at a high risk of somatic ED. This tool can facilitate entering into discussions between physicians and patients regarding erectile function.

Adult↗

Australian data and psychometric properties of the Strengths and Difficulties Questionnaire.

OBJECTIVE: We examine the Australian psychometric properties of the Strengths and Difficulties Questionnaire (SQD), a brief screening measure of behavioural and emotional problems in children and adolescents. METHOD: Using a large community sample (n = 1359) of young Australian children (4-9 years), we assessed the internal consistency, stability, and external validity of the parent-report SDQ. Normative data and cut-offs were also produced. RESULTS: Moderate to strong internal reliability was exhibited across all SDQ subscales, and support was found for the original five-factor structure of the measure. Adequate validity was evidenced in the relationship of these scales to one another, while correlations between the SDQ subscales, teacher ratings, and diagnostic interviews demonstrated sound external validity. SDQ total difficulties scores were associated with concurrent treatment status and scores over a 12-month period were stable. CONCLUSIONS: The current study of the SDQ with Australian children presents evidence of sound psychometric properties. Being the first study to empirically support the use of the SDQ in Australia, it is recommended that the youth and teacher-report forms of the measure receive similar attention in the future.

Adolescent↗

Development and validation of the Headache Needs Assessment (HANA) survey.

OBJECTIVE: To develop and validate a brief survey of migraine-related quality-of-life issues. The Headache Needs Assessment (HANA) questionnaire was designed to assess two dimensions of the chronic impact of migraine (frequency and bothersomeness). METHODS: Seven issues related to living with migraine were posed as ratings of frequency and bothersomeness. Validation studies were performed in a Web-based survey, a clinical trial responsiveness population, and a retest reliability population. Headache characteristics (eg, frequency, severity, and treatment), demographic information, and the Headache Disability Inventory were used for external validation. RESULTS: The HANA was completed in full by 994 adults in the Web survey, with a mean total score of 77.98 +/- 40.49 (range, 7 to 175). There were no floor or ceiling effects. The HANA met the standards for validity with internal consistency reliability (Cronbach alpha =.92, eigenvalue for the single factor = 4.8, and test-retest reliability = 0.77). External validity showed a high correlation between HANA and Headache Disability Inventory total scores (0.73, P<.0001), and high correlations with disease and treatment characteristics. CONCLUSIONS: These data demonstrate the psychometric properties of the HANA. The brief questionnaire may be a useful screening tool to evaluate the impact of migraine on individuals. The two-dimensional approach to patient-reported quality of life allows individuals to weight the impact of both frequency and bothersomeness of chronic migraines on multiple aspects of daily life.

Activities of Daily Living↗

Validity of the sleep subscale of the Diagnostic Assessment for the Severely Handicapped-II (DASH-II).

Currently there are no available sleep disorder measures for individuals with severe and profound intellectual disability. We, therefore, attempted to establish the external validity of the Diagnostic Assessment for the Severely Handicapped-II (DASH-II) sleep subscale by comparing daily observational sleep data with the responses of direct care staff to the sleep subscale of the DASH-II. Participants included 25 individuals with severe intellectual disability and 25 individuals with profound intellectual disability who reside in a large developmental center in central Louisiana. Four of the five items of the DASH-II were shown to have external validity. Implications of these findings for future research and practice are discussed.

Child↗

Early rheumatoid arthritis: toward tailor-made therapy.

Therapeutic possibilities for the treatment of early rheumatoid arthritis (RA) have expanded largely. New treatment modalities appear very effective with respect to relevant outcomes, such as radiographic progression. At the same time, the costs of disease-modifying antirheumatic drugs (DMARDs) have exponentially increased so that--given the rather high prevalence of RA--cost may become a limiting factor in the treatment of patients with RA. Therefore, there is a need to define the profile of those patients that should be treated with the most effective, and, unfortunately, the most costly, DMARDs. The authors describe herewith the heterogeneity of RA with respect to its most important outcomes, as well as the inability to predict those outcomes appropriately at the individual patient level. This heterogeneity of RA is not acknowledged in the modern landmark clinical trials that the authors base therapeutic decisions on, and the external validity of those trials is at stake. In this article, the authors discuss the consequences of the heterogeneity of RA in light of the perceived lack of external validity of evidence-generating landmark trials. The authors propose the following solutions to overcome this discrepancy: 1) earlier recognition of RA, and 2) appropriate prediction of treatment efficacy, because the most challenging scientific efforts may be taken in the near future in order to arrive at a tailor-made therapy for every individual presenting with RA.

Antirheumatic Agents↗

A European's perspective of COX-2 drug safety.

Regulators require less evidence to take action on drug safety concerns than would be required to grant drug efficacy claims. Such an asymmetrical view of risk points to the need for large studies to judge the safety of novel medicines compared with standard therapy as smaller sample sizes might produce unreliable and misleading toxicity signals. Large studies conducted in the classical manner are expensive and difficult to conduct. In addition, tight entry criteria and strictly protocolized care results in such trials having poor external validity. If the safety of medicines is to be judged efficiently, novel, easier to conduct, and less-expensive solutions are required. Such trials could have improved external validity were they incorporated into normal care. This paper discusses the safety of cyclooxygenase-2 inhibitors and describes the sort of studies that could be carried out to gather better safety information on their use versus standard nonsteroidal anti-inflammatory drugs.

Anti-Inflammatory Agents, Non-Steroidal↗

Clinical prediction model to characterize pulmonary nodules: validation and added value of 18F-fluorodeoxyglucose positron emission tomography.

BACKGROUND: The added value of 18F-fluorodeoxyglucose (FDG) positron emission tomography (PET) scanning as a function of pretest risk assessment in indeterminate pulmonary nodules is still unclear. OBJECTIVE: To obtain an external validation of the prediction model according to Swensen and colleagues, and to quantify the potential added value of FDG-PET scanning as a function of its operating characteristics in relation to this prediction model, in a population of patients with radiologically indeterminate pulmonary nodules. DESIGN, SETTING, AND PATIENTS: Between August 1997 and March 2001, all patients with an indeterminate solitary pulmonary nodule who had been referred for FDG-PET scanning were retrospectively identified from the database of the PET center at the VU University Medical Center. RESULTS: One hundred six patients were eligible for the study, and 61 patients (57%) proved to have malignant nodules. The goodness-of-fit statistic for the model (according to Swensen) indicated that the observed proportion of malignancies did not differ from the predicted proportion (p = 0.46). PET scan results, which were classified using the 4-point intensity scale reading, yielded an area under the evaluated receiver operating characteristic curve of 0.88 (95% confidence interval [CI], 0.77 to 0.91). The estimated difference of 0.095 (95% CI, -0.003 to 0.193) between the PET scan results classified using the 4-point intensity scale reading and the area under the curve (AUC) from the Swensen prediction was not significant (p = 0.058). The PET scan results, when added to the predicted probability calculated by the Swensen model, improves the AUC by 13.6% (95% CI, 6 to 21; p = 0.0003). CONCLUSION: The clinical prediction model of Swensen et al was proven to have external validity. However, especially in the lower range of its estimates, the model may underestimate the actual probability of malignancy. The combination of visually read FDG-PET scans and pretest factors appears to yield the best accuracy.

Aged↗

[Validity in a community (with outside verification) of primary prevention studies on hypercholesterolemia].

OBJECTIVES: The main objective of this study was to determine the degree of similarity between large primary prevention trials of hypercholesterolemia and our population of patients with dyslipidemia, in order to evaluate the external validity of these studies and their applicability to the general population. DESIGN: Descriptive retrospective study. SETTING: Tafalla Health Center in Navarra (Northern Spain), serving a population of 11 500 inhabitants.Participants. All patients older than 18 years assigned to our health center who had dyslipidemia with no antecedents of ischemic heart disease. RESULTS: The percentage of patients in our sample who satisfied the inclusion criteria used in large clinical trials ranged from 2.4% to 46%, depending on the study: AFCAPS/TexCAPS 1998, 46.2%; HPS 2002, 46.1%; WOSCOPS 1995, 10.9%; HHS 1987, 10.6%; LRC-CPPT 1984, 2.4%. CONCLUSIONS: Many of our patients (54%-97%) with dyslipidemia would not have been eligible for inclusion in earlier studies of hyperlipidemia and primary prevention. The external validity (applicability to the general population) of these studies is questionable. Decision-making in clinical practice for the primary prevention of hypercholesterolemia should be based on the risk/benefit ratio of pharmacological treatment.

Adult↗

Validation of a prediction model for the follicle-stimulating hormone response dose in women with polycystic ovary syndrome.

OBJECTIVE: To validate a published model for the prediction of the individual FSH response dose for gonadotropin induction of ovulation in polycystic ovary syndrome (PCOS). DESIGN: Structured, complete, and carefully monitored patient-based data collection to test the external validity of the prediction model. SETTING: Twenty-nine hospitals in The Netherlands. PATIENT(S): Eighty-five clomiphene citrate (CC)-resistant women with PCOS. INTERVENTION(S): Ovulation induction in a chronic low-dose step-up FSH regimen. MAIN OUTCOME MEASURE(S): Predicted individual FSH response dose, defined as follicle growth >10 mm in diameter on ultrasound. RESULT(S): The model, using the women's body mass index, CC response, initial serum FSH level, and initial serum insulin-to-glucose ratio was studied in the validation sample. Overall, the FSH response dose predicted by the model was higher than the observed response dose. The predictive performance of the model was poor, with an R(2) of 0.11, and the average prediction error was 35 IU. CONCLUSION(S): The external validity of the model predicting the individual FSH response dose was inadequate in women with CC-resistant PCOS undergoing ovulation induction with recombinant FSH in a low-dose step-up regimen.

Adult↗

Subject attrition in prevention research.

Subject attrition threatens the internal validity of substance abuse prevention studies because differences in the rate of attrition and the substance use behavior of remaining subjects in the different conditions could account for any differences found in substance use rates. Attrition threatens the external validity of prevention studies because, to the extent that study dropouts are different from remaining subjects, the results of the study may not be generalizable to study dropouts. Analysis of these threats to the validity of prevention studies should be routinely conducted. However, studies of alcohol and drug abuse prevention have generally failed to report or analyze subject attrition. Smoking prevention studies have more frequently reported attrition, and they have recently begun to analyze the degree to which attrition may affect the internal and external validity of the study. Evidence thus far suggests that differences in attrition across conditions do occur occasionally. The evidence is substantial that study dropouts are systematically more likely to smoke, to use other substances, and to score highly on other risk-taking measures.

Alcoholism↗