Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Pursuing and distancing: the construct and its measurement.

Although the concept of interpersonal pursuing and distancing has been introduced and used clinically, the lack of a reliable and valid measure has deterred its more formal investigation and relevance to personality theory. The Pursuing-Distancing (P-D) Scale, a 92-item measure of the behavioral expression of these characteristics, was developed. Evidence for internal consistency as well as internal support of construct validity is presented and discussed. An 80-item revised scale is available from the authors and is currently being evaluated in a variety of external validation settings.

Journal Article↗

Cardiac disease outcomes in clinical trials.

Randomized controlled trials (RCTs) are the gold standard for determining causality in medicine. To be considered valid, the efficacy/effectiveness of a new drug must be tested and proved through this type of scientific study. This type of trial will disclose the drug's risk profile as well as the treatment effect magnitude. The proper design, development, and analysis and presentation of results from an RCT is based on a group of well-defined methodological rules, compliance with which assures the trial's internal validity, the relative and absolute importance of the results and its applicability to populations of patients different from those included in the study sample (external validity). Among the structural and methodological components of a clinical trial--randomization, allocation concealment, confounding, similarity of study groups, measures of efficacy, statistical analysis, etc.--disease markers (endpoints, outcomes) are especially important. In the end, what an RCT is good for is to detect changes in the disease process with therapy (or preventive measures), and these changes are defined beforehand based on specific measurements--disease markers. In this paper we will present general principles for the definition of disease markers, their problems and practical use. Caution should be exercised, as this is an area of clinical epidemiology that is somewhat complex, controversial and ill-defined.

Databases, Factual↗

The Perceived Stress Questionnaire (PSQ) reconsidered: validation and reference values from different clinical and healthy adult samples.

OBJECTIVE: The aim was to translate, revise, and standardize the Perceived Stress Questionnaire (PSQ) by Levenstein et al. (1993) in German. The instrument assesses subjectively experienced stress independent of a specific and objective occasion. METHODS: Exploratory factor analyses and a revision of the scale content were carried out on a sample of 650 subjects (Psychosomatic Medicine patients, women after delivery, women after miscarriage, and students). Confirmatory analyses and examination of structural stability across subgroups were carried out on a second sample of 1,808 subjects (psychosomatic, tinnitus, inflammatory bowel disease patients, pregnant women, healthy adults) using linear structural equation modeling and multisample analyses. External validation included immunological measures in women who had suffered a miscarriage. RESULTS: Four factors (worries, tension, joy, demands) emerged, with 5 items each, as compared with the 30 items of the original PSQ. The factor structure was confirmed on the second sample. Multisample analyses yielded a fair structural stability across groups. Reliability values were satisfactory. Findings suggest that three scales represent internal stress reactions, whereas the scale "demands" relates to perceived external stressors. Significant and meaningful differences between groups indicate differential validity. A higher degree of certain immunological imbalances after miscarriage (presumably linked to pregnancy loss) was found in those women who had a higher stress score. Sensitivity to change was demonstrated in two different treatment samples. CONCLUSION: We propose the revised PSQ as a valid and economic tool for stress research. The overall score permits comparison with results from earlier studies using the original instrument.

Abortion, Spontaneous↗

Artificial intelligence-assisted histopathological diagnosis of endocervical gastric-type adenocarcinoma: a multicenter model development and validation study.

Endocervical gastric-type adenocarcinoma (GAS) is one of the most aggressive subtypes of cervical cancer and is frequently underdiagnosed due to morphological ambiguity, leading to delayed diagnosis. Despite the availability of molecular and genomic assays, their high cost, complexity, and limited reproducibility restrict clinical use. This study therefore proposes a highly sensitive artificial intelligence (AI)-assisted diagnostic system for GAS based exclusively on H&E-stained histopathological images. We included 309 slides from 96 GAS cases collected at Peking University Third Hospital from January 2018 to January 2025, representing the largest GAS cohort reported to date for AI research. In addition, we incorporated other morphologically analogous diseases, encompassing a total of 1,320 slides sourced from four categories: normal cervical mucosa (NORM), benign endocervical lesion entities (BELE), HPV-associated adenocarcinoma (HPVA), and endometrioid carcinoma with mucinous differentiation (ECMD). We developed GASPath, based on a novel multiple instance learning framework that efficiently captures fine-grained morphological variations from H&E-stained images. Beyond internal validation, GASPath was evaluated across 12 independent retrospective cohorts and further subjected to large-scale real-world validation on more than 7,000 samples from March 2024 to April 2025. Across three stages, GASPath demonstrated high performance. In internal validation (Stage I), it achieved an accuracy of 0.980 (95% CI 0.977-0.983) and an ROC-AUC of 0.995 (95% CI 0.994-0.997). In external validation (Stage II), the sensitivity reached 0.902 and improved to 0.968 with proposed strategies. For biopsy samples, GASPath achieved an ROC-AUC of 0.990 (95% CI 0.984-0.997). In large-scale real-world deployment (Stage III, n = 7,056), GASPath achieved a balanced accuracy of 0.953, with 100% sensitivity for GAS (45/45 cases correctly identified). The heatmaps highlight morphological features of GAS that are easily underestimated, such as irregular, angulated glands, subtle loss of nuclear polarity, and mild cytologic atypia, which show substantial morphological overlap with other diagnostic categories. GASPath enables high-sensitivity detection of GAS in routine H&E-stained slides, obviating the need for extensive auxiliary testing while preventing underdiagnosis and misdiagnosis. This advancement addresses a critical gap by streamlining diagnostic workflows without compromising accuracy. Its implementation could enable cost-effective, scalable AI-assisted diagnostics, potentially transforming the early detection and management of this aggressive cancer subtype.

Female↗

Condition-specific outcome measures for low back pain. Part I: validation.

A literature review of the nine most widely used, condition-specific, self-administered assessment questionnaires for low back pain has been undertaken. General and historic aspects, reliability, responsiveness and minimum clinically important difference, external validity, floor and ceiling effects and available languages were analysed for the nine most-used outcome tools. When considering which condition-specific measure to employ, the present overview on assessment tools should provide the necessary information to define the technical aspects of the nine questionnaires. These criteria, however, are only part of the consideration. In part II the construction of these scales in relationship to the measurement domains will be evaluated.

Disability Evaluation↗

TOMOCOMD-CARDD, a novel approach for computer-aided 'rational' drug design: I. Theoretical and experimental assessment of a promising method for computational screening and in silico design of new anthelmintic compounds.

In this work, the TOMOCOMD-CARDD approach has been applied to estimate the anthelmintic activity. Total and local (both atom and atom-type) quadratic indices and linear discriminant analysis were used to obtain a quantitative model that discriminates between anthelmintic and non-anthelmintic drug-like compounds. The obtained model correctly classified 90.37% of compounds in the training set. External validation processes to assess the robustness and predictive power of the obtained model were carried out. The QSAR model correctly classified 88.18% of compounds in this external prediction set. A second model was performed to outline some conclusions about the possible modes of action of anthelmintic drugs. This model permits the correct classification of 94.52% of compounds in the training set, and 80.00% of good global classification in the external prediction set. After that, the developed model was used in virtual in silico screening and several compounds from the Merck Index, Negwer's handbook and Goodman and Gilman were identified by models as anthelmintic. Finally, the experimental assay of one organic chemical (G-1) by an in vivo test coincides fairly well (100%) with model predictions. These results suggest that the proposed method will be a good tool for studying the biological properties of drug candidates during the early state of the drug-development process.

Animals↗

Prediction of hERG K+ blocking potency: application of structural knowledge.

Modelling of QT-prolongation has been performed using data for 19 structurally diverse hERG K+ channel blocking drugs taken from literature. The modelling used hydrophobicity corrected for ionisation (log D) and various 2D and 3D physico-chemical molecular descriptors. Stepwise regression produced a two parameter, interpretable and transparent QSAR with good statistical fit, including log D and the maximum diameter of molecules (Dmax). Two strategies were applied for model validation: (i) a scrambling procedure, i.e., training the total set of 19 chemicals after randomising the hERG K+ channel blocking activity data and (ii) use of external validation sets. Validation of the models showed them to be stable and statistically significant. The effect of molecular size on QT-prolongation side effect is discussed.

Anti-Arrhythmia Agents↗

Brief report: Adolescents' attitudes toward epilepsy: further validation of the Child Attitude Toward Illness Scale (CATIS).

OBJECTIVE: To examine adolescents' attitudes toward having epilepsy using the Child Attitude Toward Illness Scale (CATIS) and to provide further psychometric validation of the scale in this population. METHODS: Participants were 197 adolescents aged 11 to 17 years who completed the CATIS at two points and two external validation scales. Test-retest and internal consistency reliability and construct validity were computed. Analysis of variance was used to examine differences in attitudes according to gender, age, and epilepsy severity. RESULTS: Girls, older adolescents, and those with more severe epilepsy had more negative attitudes toward having epilepsy than boys, younger adolescents, and those with moderate or mild epilepsy, respectively. Psychometric analyses yielded excellent internal consistency reliability and good test-retest reliability. The CATIS was moderately correlated with self-esteem and mastery, supporting its construct validity. CONCLUSIONS: The CATIS is a useful and psychometrically sound tool to assess adolescents' attitudes toward having chronic illness.

Adolescent↗

The use of isometric tests of muscular function in athletic assessment.

Isometric assessment of muscular function is a popular form of testing which has been used in exercise science for over 40 years. It typically involves a maximal voluntary contraction performed at a specified joint angle against an unyielding resistance which is in series with a strain gauge, cable tensiometer, force platform or similar device whose transducer measures the applied force. Often both the maximum force and the rate of force development are recorded. These tests have generally shown high reliability in both single and multi-joint test protocols, although the maximum force is typically more reliable than rate of force development. This review outlines the reliability of isometric assessment and discusses a number of methodological considerations designed to enhance reliability and validity, including standardisation procedures, type of instructions, muscular pre-tension, testing position and joint angle. Currently, there appears to be considerable controversy as to the external validity of isometric assessment, particularly the ability of the tests to monitor changes in dynamic performance and their relationship to such performances. Indeed, a number of studies have recently shown that dynamic assessment modalities (isokinetic and isoinertial) are superior in terms of their relationship to dynamic performance and ability to discriminate between athletes of various performance levels compared with isometric assessment. This article reviews the use of isometric assessment in exercise science and consequently outlines a number of neural, mechanical and methodological factors which may have contributed to the contrasting research, and which may limit the ability of isometric assessment to relate to dynamic movement. Because of the large neural and mechanical differences between isometric and dynamic muscular actions, athletic assessment, which is dynamic in its nature, is generally most appropriately accomplished using dynamic muscular assessment methods, and in most instances isometric testing should be avoided.

Biomechanical Phenomena↗

Effectiveness of Wearable Digital Therapeutics in Improving Sleep Outcomes Among Individuals With Insomnia: Systematic Review and Meta-Analysis of Randomized Controlled Trials.

BACKGROUND: Wearable devices are increasingly used for sleep monitoring and as adjunctive treatment. Existing meta-analyses mostly pool composite digital therapies and rarely isolate stand-alone wearables or distinguish between objective and subjective end points. Whether stand-alone wearable interventions improve sleep outcomes in adults with insomnia, and which factors moderate treatment heterogeneity, remains unclear. OBJECTIVE: This study aims to evaluate the effectiveness of wearable digital interventions on sleep outcomes in adults with insomnia versus control strategies and explore moderators of effectiveness, including device-wearing position, intervention duration, and control type, using meta-regression. METHODS: This systematic review and meta-analysis was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses) 2020 statement and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta‑Analyses Literature Search Extension) guideline. Five electronic databases and clinical trial registries were searched from inception to May 18, 2026. Eligible studies were randomized controlled trials (RCTs) evaluating wearable digital interventions in adults with insomnia compared with sham, waitlist, usual care, or active control conditions and had an intervention duration of at least 1 week. Study screening, data extraction, and risk-of-bias assessment were carried out independently by 2 reviewers. Pooled estimates were calculated using a restricted maximum likelihood random-effects model with the Hartung-Knapp-Sidik-Jonkman correction. Heterogeneity was assessed using the I² statistic, and 95% prediction intervals (PIs) were calculated for the primary analyses. The certainty of evidence was rated using the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) approach. RESULTS: Sixteen RCTs (N=910) were included. Wearable digital interventions were associated with a significant reduction in objective sleep-onset latency (SOL; mean difference [MD] -4.52, 95% CI -8.38 to -0.67, PI -9.52 to 0.47 min) and a significant improvement in subjective sleep efficiency (SE; MD 2.00%, 95% CI 1.90%-2.11%, PI 1.85%-2.15%). Subjective total sleep time (TST) also showed a significant increase (MD 19.11, 95% CI 2.98-35.24, PI -16.20 to 54.43 minutes). Meta-regression showed that control type, intervention duration, and device location did not explain the heterogeneity of the insomnia severity index (ISI) (R²=0). Sensitivity analysis confirmed the robustness of pooled ISI estimates, and an Egger test indicated no small-study effects (P=.07). Certainty of evidence ranged from moderate to high. CONCLUSIONS: Wearable digital interventions provide selective benefits for objective SOL, subjective SE, and subjective TST in adults with insomnia, with no improvement in overall ISI. Despite statistically significant effects on several sleep parameters, wide PIs, substantial heterogeneity, and limited study numbers indicate preliminary, nonconclusive findings. Wearables should be viewed as affordable adjunctive tools requiring further validation, not substitutes for first-line cognitive behavioral therapy for insomnia. Large-scale, long-term RCTs with standardized protocols and patient-level external validation are required to consolidate the evidence base.

Humans↗

Correlates of the impostor phenomenon among undergraduate entrepreneurs.

The impostor phenomenon describes the self-attribution of success to luck and interpersonal skills rather than to intelligence and ability, despite external validation to the contrary. Evidence suggests the presence of impostor characteristics among a group of 63 undergraduate entrepreneurs. More intense impostor feelings were associated with an external locus of control and a stronger perceived effect of work on family life. Implications for entrepreneurial performance are discussed and questions for research are presented.

Adult↗

Screening for personality disorders in a nonclinical population.

Based on the Inventory of Interpersonal Problems (IIP), the IIP-PD and the IIP-C screening scales were developed to distinguish personality disorder (PD) from non-PD and Cluster C from other PD, respectively, in a clinic population. Two studies were conducted to determine (a) validity and reliability of these IIP scales for PD screening in a nonclinical population, (b) specificity of IIP-C for identifying Cluster C, and (c) usefulness of the IIP scales for screening Cluster A. College students were screened using the IIP scales (Study 1, N = 454, Study 2, N = 87). High and low scorers completed PD-related questionnaires in Study 1 and a clinical interview for PD symptomatology in Study 2. Results indicated strong test-retest reliability, internal consistency, and factorial, convergent, and external validity. The scales tapped a common deficit in interpersonal relatedness, with some distinction between externalizing and internalizing dimensions, respectively, and both scales were positively and significantly associated with schizotypal traits. In conclusion, the IIP-PD and IIP-C are useful and valid screening instruments for identifying any versus no PD in nonclinical populations.

Adolescent↗

Development of a rating scale for quantitative measurement of the alcohol withdrawal syndrome.

The alcohol withdrawal syndrome consists of autonomic, neurological and mental symptoms. For its assessment, these symptoms have to be rated in a quantitative and valid manner. We developed a new rating scale for mild and moderate alcohol withdrawal states. Difficulty, discrimination coefficient, internal consistency, and the principal component analysis were assessed. External validation was tested on a separate sample of inpatients. Eight of 12 original items fulfilled test-theoretical criteria. From these a psychosensory and an autonomic factor have been extracted. This instrument can be used repeatedly for clinical assessment as well as for evaluation of the alcohol withdrawal syndrome in clinical drug studies.

Adult↗

A method for scoring the pain map of the McGill Pain Questionnaire for use in epidemiologic studies.

Identifying and quantifying the location of pain may be important for understanding specific functional impairments in elderly populations. The purpose of the present analysis was two-fold: first, to describe the reliability of a scoring method for the McGill Pain Map (MPM), and second, to validate the method of scoring the MPM as a tool for assessing areas of body pain in an epidemiologic study. In interviews performed at the subjects' homes, 411 community dwelling Mexican-American and non-Hispanic white subjects aged 65-74 from the San Antonio Longitudinal Study of Aging (SALSA) were asked to describe the location of their pain on the map of the human body included in the McGill Pain Questionnaire. The location of pain was scored by overlaying the survey figures with a MPM template divided into 36 anatomical areas. Inter- and intra-rater agreement among three raters was measured by calculating a kappa statistic for each of the body areas, and an intraclass correlation coefficient for the total number of painful areas (NPA). Internal validity was measured by Spearman's rho between the NPA and the Present Pain Index (PPI) and Pain Rating Index (PRI) of the McGill Pain Questionnaire, and external validity by correlation between NPA and the Perceived Health (PH), Amount of Bodily Pain (APB), and Pain Interference with Work (PIW) items of the Medical Outcomes Study, and the Perceived Physical Health (PPH) question of the San Antonio Heart Study. Average inter-rater agreement for individual MPM areas was 0.92 +/- 0.01, and average agreement for NPA was 0.96 +/- 0.01. Intra-rater agreement for individual areas averaged 0.94 +/- 0.01, and for NPA = 0.99 +/- 0.001. Pain in one or more areas was present in 47.7% of the subjects. For the whole sample, correlations between NPA and the validation indices were: PPI (0.91), PRI (0.89), PH (0.25), ABP (0.64), PIW (0.49), and PPH (0.20). Among the 196 subjects with pain, correlations were: PPI (0.34), PRI (0.34), PH (0.19), ABP (0.21), PIW (0.38), and PPH (0.19)-p < 0.01 for all correlations. In conclusion, we have developed a reliable method of scoring the MPM and have shown evidence of its validity in a community-based sample of elderly subjects. Patterns of painful body areas may be associated with specific diseases and functional impairments.

Aged↗

Developmental trajectories of offending: validation and prediction to young adult alcohol use, drug use, and depressive symptoms.

This longitudinal study extended previous work of Wiesner and Capaldi by examining the validity of differing offending pathways and the prediction from the pathways to substance use and depressive symptoms for 204 young men. Findings from this study indicated good external validity of the offending trajectories. Further, substance use and depressive symptoms in young adulthood (i.e., ages 23-24 through 25-26 years) varied depending on different trajectories of offending from early adolescence to young adulthood (i.e., ages 12-13 through 23-24 years), even after controlling for antisocial propensity, parental criminality, demographic factors, and prior levels of each outcome. Specifically, chronic high-level offenders had higher levels of depressive symptoms and engaged more often in drug use compared with very rare, decreasing low-level, and decreasing high-level offenders. Chronic low-level offenders, in contrast, displayed fewer systematic differences compared with the two decreasing offender groups and the chronic high-level offenders. The findings supported the contention that varying courses of offending may have plausible causal effects on young adult outcomes beyond the effects of an underlying propensity for crime.

Adolescent↗

The application of special technologies in diagnostic anatomic pathology: is it consistent with the principles of evidence-based medicine?

Proponents of evidence-based medicine (EBM) have emphasized the need to consider the quality of different sources of medical information and have proposed various methods to integrate available "best evidence" into rules, guidelines and other diagnostic, therapeutic and prognostic models. The various factors that can affect the internal validity of studies in anatomic pathology, such as interobserver variability, use of retrospective rather than prospective data and others, are reviewed. The need for testing for the external validity of the results of anatomic pathology studies is introduced, using "test sets" of cases that have not been used to generate the classification or prognostic models. This methodology has been seldom used in anatomic pathology to validate the generalizability of various "entities," usefulness of diagnostic tests under different conditions and other information. Basic concepts of meta-analysis for research synthesis are introduced; these methods have been seldom used in anatomic pathology to integrate information from different studies using quantitative techniques rather than summary tables that merely list the results of various publications. The potential use of decision analysis and value of information analysis for the adoption of new tests is briefly discussed.

Data Interpretation, Statistical↗

Neurobehavioral assessment of outcome following traumatic brain injury in rats: an evaluation of selected measures.

Neurobehavioral assessment of outcome has played an integral part in traumatic brain injury (TBI) research. Given the fundamental role of neurobehavioral measurement, it is critical that the tasks used are of the highest psychometric quality. The purpose of this paper is to evaluate several, commonly used neurobehavioral measures along the dimensions of reliability, sensitivity, and validity. Using both the midline and lateral fluid-percussion injury models, nine neurobehavioral measures were evaluated that assessed three different neurobehavioral constructs. Reflex suppression was measured by the duration of the suppression of the pinna, corneal, and righting reflexes. Vestibulomotor function was assessed with the beam-balance, beam-walking, and rotorod tasks. Cognitive function was evaluated by three measures of Morris water maze performance (goal latency, path length, cumulative distance). The evaluation of the reliability of the nine neurobehavioral measures found that all had acceptably high reliability coefficients (0.79 or higher). The analysis of each measure's sensitivity to injury found that all measures were capable of detecting injury-induced impairments. However, there were some substantial differences in the sensitivity of the measures of vestibulomotor and maze performance: the rotorod was the most sensitive vestibulomotor measure and goal latency and path length were equally sensitive measures of maze performance. In the assessment of validity, the results of a factor analysis supported the convergent and discriminative validity of the measures. And in cases in which the preclinical and clinical research have assessed the same construct, the animal model neurobehavioral measures had predictive (or external) validity. Thus, according to the psychometric standards by which measurement instruments are evaluated, the results indicated that these measures provide a valid assessment of neurobehavioral function after fluid percussion TBI.

Animals↗

Effects of reward and familiarity of reward agent on spontaneous play in preschoolers: a field study.

A field experiment was conducted with preschool children to test the effect of rewards on a familiar, spontaneous play activity, in conditions as close as possible to the children's natural school context, and to examine the role of familiarity of the person who administered rewards. In three experimental conditions, children were rewarded either by their own teacher or by an unknown adult for playing with toys at the school playground and stayed with either the teacher or the unknown adult in the remaining part of reward sessions. Spontaneous play was significantly reduced by the reward relative to baseline levels and recovered after a 3-wk. interval. However, no difference due to familiarity/unfamiliarity of reward agent could be found. Results, discussed in terms of an incentive contrast hypothesis, attest to the generality and external validity of the undermining effects of rewards.

Child, Preschool↗