Search PubMedSearch

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Internal and external validity in two studies that compared treatment methods.

Research comparing the effectiveness of two treatments offers both strengths and weaknesses for occupational therapy. Although it is worthwhile to determine which of two treatments works best for a particular problem, methodological problems may arise that preclude a valid conclusion. To draw valid conclusions from research, criteria for internal validity and external validity must be satisfied. The two preceding articles in this issue are examples of studies that used between-groups experimental methodology to compare the effectiveness of two different treatments. This paper evaluates the above-mentioned studies on the basis of principles of internal and external validity. One of these studies (Jongbloed, Stacey, & Brighton, 1989) was truly experimental, whereas the other study (Groves & Rider, 1989) was quasi-experimental. Results from both studies were similar because, in each study, both of the treatment groups improved, but there were no significant differences between treatments. Absence of a true control group in both studies presented limitations on the conclusion that both treatments worked equally well.

Analysis of Variance

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75 161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et al., Nanda et al., Naylor et al., and Van Leeuwen et al., each showing fair discrimination. The Teede et al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et al. and van Leeuwen et al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

External validity in the assessment of intellectual development in adulthood.

The relation of intelligence and competence is discussed and external validity issues are examined for the dimensions of settings, measurement variables, treatment variables and experimental units. It is argued that external validity across situations and life stages cannot be obtained for any single measure of intellectual ability. External validity problems are exacerbated beyond young adulthood since single criterion goals comparable to that of educational aptitude in work with the young are not available, and tasks do not retain ecological validity when the situational context of the individual under study changes due to developmental progression and idiosyncratic modification of individual life situation and roles. External validity in adulthood must therefore be addressed by examining task-by-person-by-situation interfaces separately for different life stages and across cohort groupings. A major test construction and validation program is outlined, and examples are given showing how some of the aspects of such a program can be operationalized.

Cognition

External validation, repeat determination, and precision of risk estimation in misclassified exposure data in epidemiology.

STUDY OBJECTIVE: The aim was to quantify the difference in precision of risk estimates in epidemiology between the situations where misclassification of exposure is corrected for by external validation and where it is corrected for by internal repeat measurement. Precision was measured in terms of the expected width of the 95% confidence interval on the odds ratio. DESIGN: In a hypothetical case-control study, first with 100 cases and 100 controls, then with 100 cases and 1000 controls (the latter to approximate the cohort study situation), expected estimated odds ratios and confidence intervals were calculated based on postulated underlying true odds ratios and misclassification error rates. The sizes of the confidence intervals using the two design strategies were compared, based on the same number of subjects receiving internal repeat measurements as were used in the external validation study. MAIN RESULTS: Confidence intervals obtained using internal repeat measurement were considerably narrower than those using external validation. Both methods yielded approximately correct point estimates. CONCLUSIONS: In terms of precision, it is preferable to correct for misclassification using internal repeat measurement rather than external validation.

Analysis of Variance

An external validity study of the MMPI Personality Disorder Scales.

This study examined the external validity of the MMPI Personality Disorder (PD) scales. Patients diagnosed with personality disorders (n = 23), according to a structured interview, were contrasted with medical control subjects and psychiatric patients with no personality disorder (n = 33). The MMPI PD scales could discriminate patients with any personality disorder and patients with specific clusters of personality disorders. Convergent validity was demonstrated by high correlations between the MMPI PD scales and the comparable Millon Clinical Multiaxial Inventory scales, as well as the number of DSM-III personality disorder symptoms. This study represents a preliminary step in the external validation of the MMPI PD scales.

Adult

The external validity of age- versus IQ-discrepancy definitions of reading disability: lessons from a twin study.

Recent research has raised the question of whether age- and IQ-discrepancy forms of reading disability (RD) are distinguishable in terms of either their underlying linguistic deficit or their response to treatment, thus threatening the external validity of the traditional distinction between specific reading retardation and reading backwardness. The present study pursued the external validity of this distinction in three domains: (a) genetic etiology, (b) sex ratio and clinical correlates, and (c) neuropsychological profiles. Each of these domains was explored in the RD (n = 640) and control (n = 436) twins participating in the Colorado Reading Project (514 males, 562 females, with an overall mean age of 12.42 years). Little evidence for external validity was found in terms of the clinical correlates of attention deficit-hyperactivity disorder (ADHD), immune disorders, or handedness. Most importantly, there was no evidence of differential genetic etiology of the two phenotypes in this sample, in that deficits in both phenotypes were similarly heritable (h2g = .40 and .46 for age and IQ phenotypes, respectively) and the genetic correlation between them was high (rG = .88 to .96). However, the genetic and neuropsychological profile analyses did suggest that age- and IQ-discrepancy definitions of RD may relate differentially to component reading processes, such as phonological awareness and orthographic coding.

Attention Deficit Disorder with Hyperactivity

Improving depression severity assessment--II. Content, concurrent and external validity of three observer depression scales.

The Hamilton Depression Scale (HAMD), the Montgomery-Asberg Depression Rating Scale (MADRS) and the Bech-Rafaelsen Melancholia Scale (BRMS) were compared with respect to content, concurrent and external validity in sample of 130 patients with a major depressive episode. The three scales did equally well in concurrent and external validity. The HAMD showed some deficiencies in content validity. The consequences for depression severity assessment are discussed.

Adult

Panel attrition and external validity in adolescent substance use research.

Panel attrition threatens external validity in adolescent substance use research. A 7-year adolescent panel was examined to determine whether attrition effects varied by (a) type of substance assessed and (b) method of measurement and type of statistical analysis. Chi-squares and multivariate analyses of variance revealed that study dropouts were more likely to use substances and reported higher mean use of substances at baseline than stayers; attrition effects varied by substance; and mean use comparisons were more likely to detect attrition effects than use-nonuse comparisons. Implications of these findings for adolescent substance use research are discussed.

Adolescent

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure

Do smoking prevention programs really work? Attrition and the internal and external validity of an evaluation of a refusal skills training program.

This study investigated the effects of a smoking prevention program that emphasized refusal skills training on 1730 adolescents in three high schools and six middle schools. Classes within these schools were randomly assigned to treatment or no-treatment conditions to avoid confounding schools with treatment condition. The effects of attrition on the internal and external validity of the study were examined. Although the results indicated an apparent effect of the program at the 1-year follow-up in deterring continued smoking among those who were smoking at pretest, this result may have been due to a higher rate of attrition among high-rate smokers in the treatment condition than in the control condition. Attrition also affected external validity. Across both conditions, subjects who were smoking at pretest and who were at risk to smoke were more likely to be missing at follow-up. The program did have an effect on the refusal skills of participants and the validity of this effect was not jeopardized by differential attrition.

Adolescent

A methodological approach to enhance external validity in simulation based research.

Simulation methodology, as exemplified by use of vignettes, has many advantages, such as standardization of data collection procedure, control of extraneous variables, and manipulation of variables of interest. The main shortcoming is its artificiality and therefore its limited external validity. To determine the degree to which findings from simulation research are transferable to the real world, a comparison can be made using the same measure on artificial and real situations, or the researcher can determine the degree to which a score on the simulation compares with a score in the field. The predictive validity approach was used in a study that used an artificial method to study patient assault. In this study, the intent was to determine the degree of accuracy that the subjects' causal attribution scores on an assault vignette were predictive of causal attribution scores in the actual assault situation.

Guilt

Testing the internal and external validity of a simplified dental caries index on an adult population.

This analysis of a caries index, proposed in 1966 to WHO as a simplified method of measurement, as tested on a 16- to 45-year-old population who were seeking dental care at the University of Minnesota School of Dentistry revealed several weaknesses associated with the index. An analysis of the external validity of this index, a comparison with subjects' DMFS scores, revealed a correlation coefficient of 0.71. Although the index purports to measure the prevalence and severity of dental caries by dividing the dentition into five zones representing increasing severity of dental caries experience, an analysis of this index's internal validity, i.e. whether these five zones truly represent a rank-order scale of severity, revealed misclassification rates of from 21% for the total population up to 44% for a subgroup. When zones were recombined to reduce the misclassification rates, the descriptive capabilities of the index were greatly reduced as most subjects were then classified in only one or two of the zones.

Adolescent

Intracerebral hemorrhage: external validation and extension of a model for prediction of 30-day survival.

We report validation of a previously reported logistic regression model for predicting 30-day survival after supratentorial intracerebral hemorrhage using independent, prospectively collected data. The original model, using initial Glasgow Coma Scale score, hemorrhage size, and pulse pressure, accounted for mortality or survival at 30 days in 92% of patients in the Pilot Stroke Data Bank with a sensitivity of 0.84 and a specificity of 0.96. For external validation, the model was used to predict 30-day status for each patient in the Main Phase Stroke Data Bank for whom complete risk factor information was available. Overall, 90% of patients' outcomes were correctly predicted with a sensitivity of 0.85 and a specificity of 0.92. Two factors not collected in the Pilot Stroke Data Bank, hyperglycemia and intraventricular hemorrhage extension, were assessed to determine if they provided additional predictive information on 30-day mortality. Intraventricular hemorrhage extension contributed significant predictive information in a logistic regression, whereas hyperglycemia did not. The resulting four-factor model with an interaction term (intraventricular hemorrhage extension and Glasgow Coma Scale score) correctly classified the survival status of 94% of patients at 30 days. A more general outcome, death or failure to achieve a "good" Activities of Daily Living Score by one year, was analyzed with respect to the same four factors. The resulting model correctly classified 95% of the patients in the cohort.

Blood Glucose

Identification and external validation of a prognostic signature based on myeloid-derived suppressor cells-related LncRNAs to evaluate survival prognosis and treatment efficacy in invasive breast carcinoma.

BACKGROUND: Originating in the hematopoietic tissue, myeloid-derived suppressor cells (MDSCs) significantly contribute to tumor-related immunological processes. However, their relationship with long noncoding RNAs (lncRNAs) and breast cancer remains incompletely understood. In this study, we introduced MDSCs-associated lncRNAs as novel prognostic biomarkers to assess outcomes in patients with invasive breast carcinoma (BRCA). METHODS: Information regarding BRCA cases, including clinical and genomic details, was obtained from the TCGA repository. Predictive indicators were discovered, and their reliability underwent thorough verification. A clinically useful nomogram was developed following application-based validation. Additional investigations encompassed functional analysis, TMB assessment, TME profiling, immunotherapy efficacy forecasting, and drug sensitivity testing along with target identification. Long non-coding RNA expression was measured using reverse transcription quantitative PCR. RESULTS: A risk stratification model incorporating eight MDSCs-related lncRNAs effectively predicted patient outcomes. Kaplan-Meier (K-M) survival analysis clearly indicated a much worse prognosis among patients classified as high-risk (p&#xa0;<&#xa0;0.001). The nomogram accurately forecasted overall survival (OS). Analysis of functional enrichment revealed that pathways associated with epithelial cells showed activity among patients at higher risk. Characterization of the tumor microenvironment showed increased immune cell presence in those classified as low-risk. Conversely, individuals with greater risk displayed higher tumor mutational burden. TIDE and IPS analyses indicated superior immunotherapy responsiveness in the low-risk BRCA subgroup. Among 47 drugs with notable IC50 variations, Ribociclib, PD173074, KU-55933, NU7441, and nutlin-3a exhibited lower IC50 values within the low-risk group, whereas Lapatinib demonstrated greater efficacy among the high-risk group. Moreover, 10 potential therapeutic agents and their targets were predicted for high-risk patients. RT-qPCR validation confirmed the robustness of the model. CONCLUSIONS: We successfully verified a new model of molecular markers of MDSCs-related lncRNAs, offering critical insights for predicting outcomes and guiding therapeutic decisions in BRCA cases.

Bioinformatics

External validity of the new Devereux Adolescent Behavior Rating scales.

The discriminant and concurrent validity of the five new scales for the Devereux Adolescent Behavior Rating Scale (DAB) was explored using a heterogeneous sample of psychiatric and substance abuse patients. Consistent with predictions, the substance abuse patients scored higher on the Acting Out Behaviors (AOB) and Heterosexual Interests (HI) scales, and psychiatric patients scored higher on the Psychotic Behaviors scale. Gender differences also were found, including boys being rated higher on Acting Out Behaviors, and girls higher on Heterosexual Interests. The new DAB scales demonstrated sufficient concurrent validity using a thorough record review and a patient rating scale (the Child Behavior Checklist [CBCL]). The Neurotic/Dependent Behaviors scale (NDB) showed a consistent relationship with substance abuse and several other measures of more externalizing behaviors, in addition to the predicted relationships with anxious, tense, and dependent behaviors. The Withdrawn/Timid Behaviors scale (WTB) proved to be a purer measure of internalizing behaviors in both sexes.

Acting Out