Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Subjective experience and related symptoms in schizophrenia.

We had previously extracted two types of subjective experience of schizophrenia (SES); first, a feeling of inadequacy in stream of speech, thought, and action, associated with a distorted sense of self, and second, a feeling that excessive thoughts are filling and sticking to one's head, causing negative affective burden such as misery and oppression. This study tried to validate their content using conventional symptom clusters as external validators. Subjects were 63 patients from two hospitals in Tokyo meeting ICD-10 criteria for schizophrenia. Positive, negative, and depressive psychopathology were measured by the Brief Psychiatric Rating Scale (BPRS), Scale for the Assessment of the Negative Symptoms (SANS), and Hamilton's Depression Scale (HDS). The two types of SES were measured by an original scale. The first type of SES correlated significantly with the negative symptoms of alogia, avolition, and attention, whereas the second correlated with positive and depressive symptoms. To analyze how schizophrenia is experienced by patients, qualitative and comprehensive descriptions, such as indicated by our subjective factors, will be useful.

Adult↗

Dimensionality, responsiveness and standardization of the Bech-Rafaelsen Mania Scale in the ultra-short therapy with antipsychotics in patients with severe manic episodes.

OBJECTIVE: Typical antipsychotics have their indication in the ultra-short (first week) treatment of severe episodes of mania. In this setting the Bech-Rafaelsen Mania Scale (MAS) was psychometrically compared with the Clinical Global Impression scale (CGI) to assess its ability to measure response. METHOD: Ratings on patients with marked to severe mania (n = 80) who participated in the clinical trials to evaluate the ultra-short antimanic effect of zuclopenthixol acetate were assessed. The MAS was analysed for internal validity (total score a sufficient statistic) and for external validity. RESULTS: The MAS was shown to have a high internal validity showing onset of action already after days of treatment. After 6 days of treatment 53% of the patients responded according to the MAS but only 30% according to the CGI. The difference was statistically significant. CONCLUSION: The MAS has been found to be a valid scale to measure early onset of action and response in the ultra-short antimanic treatment with typical antipsychotics.

Aged↗

An individualized nomogram for predicting progression-free survival in systemic anaplastic large cell lymphoma: a multicenter, retrospective, and internally validated study.

OBJECTIVES: To develop an individualized nomogram for predicting disease progression risk in systemic anaplastic large cell lymphoma (sALCL). METHODS: Independent predictors of progression-free survival (PFS) were identified using Cox regression in a multicenter retrospective cohort of 109 sALCL patients (2010-2022). These were incorporated into a three-factor nomogram, evaluated via bootstrapped internal validation (1000 resamples), ROC analysis, C-index, decision curve analysis (DCA), and clinical impact curve (CIC). RESULTS: A total of 29 PFS events occurred during a median follow-up of 31 months. Multivariable modelling selected serum β2-microglobulin elevation, extranodal disease, and front-line chemotherapy choice (CHOP versus CHOPE or BV+CHP) as autonomous progression drivers. Upon internal bootstrap validation, the nomogram yielded strong prognostic accuracy, achieving AUCs of 0.81, 0.85 and 0.87 for 1-, 3- and 5-year progression-free survival, alongside a corrected C-index of 0.779 (95% CI: 0.699 - 0.861). Calibration plots showed close agreement between predicted and observed outcomes, while DCA confirmed superior net clinical benefit versus conventional IPI or Ann Arbor stratification across multiple decision thresholds. CONCLUSION: This first sALCL-specific nomogram integrates clinical and treatment variables to provide personalized PFS risk estimation. While internally validated, this exploratory, observation-based tool requires external validation and recalibration in prospective cohorts before clinical implementation.

Humans↗

Short-term effects of ambient oxidant exposure on mortality: a combined analysis within the APHEA project. Air Pollution and Health: a European Approach.

The Air Pollution and Health: a European Approach (APHEA) project is a coordinated study of the short-term effects of air pollution on mortality and hospital admissions using data from 15 European cities, with a wide range of geographic, sociodemographic, climatic, and air quality patterns. The objective of this paper is to summarize the results of the short-term effects of ambient oxidants on daily deaths from all causes (excluding accidents). Within the APHEA project, six cities spanning Central and Western Europe provided data on daily deaths and NO2 and/or O3 levels. The data were analyzed by each center separately following a standardized methodology to ensure comparability of results. Poisson autoregressive models allowing for overdispersion were fitted. Fixed effects models were used to pool the individual regression coefficients when there was no evidence of heterogeneity among the cities and random effects models otherwise. Factors possibly correlated with heterogeneity were also investigated. Significant positive associations were found between daily deaths and both NO2 and O3. Increases of 50 micrograms/m3 in NO2 (1-hour maximum) or O3 (1-hour maximum) were associated with a 1.3% (95% confidence interval 0.9-1.8) and 2.9% (95% confidence interval 1.0-4.9) increase in the daily number of deaths, respectively. Stratified analysis of NO2 effects by low and high levels of black smoke or O3 showed no significant evidence for an interaction within each city. However, there was a tendency for larger effects of NO2 in cities with higher levels of black smoke. The pooled estimate for the O3 effect was only slightly reduced, whereas the one for NO2 was almost halved (although it remained significant) when two pollutant models including black smoke were applied. The internal validity (consistency across cities) as well as the external validity (similarities with other published studies) of our results on the O3 effect support the hypothesis of a causal relation between O3 and all cause daily mortality. However, the short-term effects of NO2 on mortality may be confounded by other vehicle-derived pollutants. Thus, the issue of independent NO2 effects requires additional investigation.

Europe↗

Circular instead of hierarchical: methodological principles for the evaluation of complex interventions.

BACKGROUND: The reasoning behind evaluating medical interventions is that a hierarchy of methods exists which successively produce improved and therefore more rigorous evidence based medicine upon which to make clinical decisions. At the foundation of this hierarchy are case studies, retrospective and prospective case series, followed by cohort studies with historical and concomitant non-randomized controls. Open-label randomized controlled studies (RCTs), and finally blinded, placebo-controlled RCTs, which offer most internal validity are considered the most reliable evidence. Rigorous RCTs remove bias. Evidence from RCTs forms the basis of meta-analyses and systematic reviews. This hierarchy, founded on a pharmacological model of therapy, is generalized to other interventions which may be complex and non-pharmacological (healing, acupuncture and surgery). DISCUSSION: The hierarchical model is valid for limited questions of efficacy, for instance for regulatory purposes and newly devised products and pharmacological preparations. It is inadequate for the evaluation of complex interventions such as physiotherapy, surgery and complementary and alternative medicine (CAM). This has to do with the essential tension between internal validity (rigor and the removal of bias) and external validity (generalizability). SUMMARY: Instead of an Evidence Hierarchy, we propose a Circular Model. This would imply a multiplicity of methods, using different designs, counterbalancing their individual strengths and weaknesses to arrive at pragmatic but equally rigorous evidence which would provide significant assistance in clinical and health systems innovation. Such evidence would better inform national health care technology assessment agencies and promote evidence based health reform.

Evidence-Based Medicine↗

Teaching undergraduate nursing research: a narrative review of evaluation studies and a typology for further research.

Before knowledge about strategies for teaching undergraduate nursing research can be synthesized, it must be demonstrated that studies evaluating teaching outcomes are valid. A narrative review using Cook and Campbell's validity framework for experimental and quasi-experimental studies was done to appraise the validity of the eight studies in which teaching strategies for undergraduate research have been evaluated. Although the studies sampled had evidence of uncontrolled threats to validity, the research reports emphasized the teaching strategy's "contextual validity" rather than the study's internal or external validity. This insight stimulated questions about the appropriateness of a traditional perspective on validity for research designed to evaluate teaching strategies. Accordingly, a "typology for research on strategies for teaching research," loosely based on Beer and Bloomer's schema for evaluation research, was proposed. The typology consists of four categories of evaluation studies that vary according to purpose, methodological assumptions, types of validity emphasized, and student sample size. The four categories, which are minimally represented in the current literature, are complementary approaches for further evaluation studies of teaching strategies in undergraduate nursing research.

Data Interpretation, Statistical↗

A quality assessment of randomized control trials of primary treatment of breast cancer.

The methodology of randomized control trials (RCTs) of the primary treatment of early breast cancer has been reviewed using a quantitative method. Sixty-three RCTs comparing various treatment modalities tested on over 34,000 patients and reported in 119 papers were evaluated according to a standardized scoring system. A percentage score was developed to assess the internal validity of a study (referring to the quality of its design and execution) and its external validity (referring to presentation of information required to determine its generalizability). An overall score was also calculated as the combination of the two. The mean overall score for the 63 RCTs was 50% (95% confidence interval [CI] = 46% to 54%) with small and nonstatistically significant differences between types of trial. The most common methodologic deficiencies encountered in these studies were related to the randomization process (only 27 of the 63 RCTs adopted a truly blinded procedure), the handling of withdrawals (only 26 RCTs included all patients in the analyses), the description of the follow-up schedule (only 12 RCTs reported adequately), the report of side effects (adequate information given in 33 RCTs), and the description of the patient population (satisfactory in 29 RCTs). Telephone calls to the principal investigators improved the quality scores by seven points on a scale of 100, indicating that some of the deficiencies lay in reporting rather than performance. There was evidence that quality has improved over time and that the increasing tendency of involving a biostatistician in the research team was positively associated with the improvement of the internal validity but not with the external.

Breast Neoplasms↗

Consensus kNN QSAR: a versatile method for predicting the estrogenic activity of organic compounds in silico. A comparative study with five estrogen receptors and a large, diverse set of ligands.

Quantitative structure-activity relationships (QSARs) have proved increasingly useful for predicting the biological activities of molecules (e.g., their binding affinities to different receptors) and can be used in environmental chemistry as a preliminary tool for screening the activities of untested molecules, producing valuable information on which compounds should be tested more thoroughly with experimental affinity assays or in animals. The predictive ability of the consensus kNN QSAR method is corroborated here using a diverse set of 245 compounds, which have been assayed for their relative binding affinities to the estrogen receptor of four species: human (ER alpha and ER beta), calf, mouse, and rat. Leave-one-out cross-validation (LOO-CV) and gamma-randomization tests were applied to the QSAR models for internal validation, and separate training and test sets were used for external validation. The internal predictive abilities of the consensus models for all five data sets were convincing, with cross-validated correlation coefficients (LOO-CV q2 values) varying from 0.69 (human ER beta data) to 0.79 (human ER alpha data). The external predictive abilities were also encouraging, as the predictive r2 scores (pr-r2 values) varied from 0.62 (human ER beta data) to 0.77 (calf and mouse data). The results indicate that consensus kNN QSAR is a feasible method for rapid screening of the estrogenic activity of organic compounds.

Animals↗

Deep Learning on Histologic Slides Accurately Predicts Consensus Molecular Subtypes and Spatial Heterogeneity in Colon Cancer.

Colon cancer (CC) is the third most prevalent cancer type. It is highly heterogeneous, particularly in terms of molecular profiles, which have both prognostic and predictive impacts on the treatment efficacy. However, CC treatment in adjuvant situations is currently guided solely by T and N staging. In this context, consensus molecular subtypes (CMSs) were introduced to stratify patients with CC based on molecular profiles. Recent studies have shown that CMS can be heterogeneous in CC, leading to a worse prognosis. This study focused on predicting CMS and its heterogeneity in CC using deep learning on digitized hematoxylin and eosin ± saffron-stained whole-slide images. Data and whole-slide images of 1996 patients from the PETACC-8, The Cancer Genome Atlas-COAD, and PRODIGE-13 cohorts were used. The model is trained to predict a 4-dimensional CMS vector, reflecting intratumor heterogeneity (ITH). It comprises a self-supervised model for embedding image patches into vectors and a weakly supervised model predicting CMS calls. Ground-truth CMS scores are obtained with the CMSclassifier package. Interpretability analyses are performed at the slide and patch levels. For homogeneous tumors, the model trained on PETACC-8 achieves 93.0% (±1.4%) macroaverage area under the curve in internal cross-validation and 94.4% macroaverage area under the curve in external validation over PRODIGE-13, whereas the The Cancer Genome Atlas-COAD model reaches 85.4% (±3.0%) in cross-validation and 92.4% over PRODIGE-13. The trained models also provide spatial distributions of CMS across tumor slides and associate specific histologic features with each CMS. Finally, the models are able to predict ITH. The results show that a deep learning model trained on routine histology slides is capable of providing an efficient and robust method for predicting CMS and characterizing a patient's ITH, paving the way for the routine consideration of CMS/ITH in clinical decision making in the adjuvant setting.

Humans↗

Evaluation of community-based injury prevention programmes: methodological issues and challenges.

The evaluation of comprehensive community-based injury prevention programmes is complex and poses many methodological challenges. There is little consensus in contemporary literature about the most appropriate methods of evaluating these programmes. This study employed a systematic literature review to examine evaluations of 16 community-based injury prevention programmes with regard to key methodological issues and challenges. Three aspects of the evaluated programmes were analysed: assessed elements (context, structure, process, impact, and outcome); study design; and methodological issues addressed. The results showed that context, structure and process assessments were the most neglected aspects of the evaluation studies. The programmes were typically described with minimal discussion of how the context may have influenced the effectiveness. The process (activities) was described rather than evaluated against appropriate standards of comparisons. Impact evaluations adhered more closely to documented guidelines, but half of the evaluations did not include impact variables. Outcome evaluations focused on injury incidence. Most evaluations employed some qualitative methods, but the vast majority of methods used were quantitative. This study indicated that the quasi-experimental study design has become an accepted norm for the evaluation of community-based injury prevention programmes. Most of the evaluations contained explicit details of the methodology used and of the choices related to the methodology. While threats to internal validity were identified in most studies, problems related to external validity and construct validity were largely overlooked by the evaluators.

Community Health Planning↗

Subtypes of aggression and their relevance to child psychiatry.

OBJECTIVE: To review the evidence for qualitatively distinct subtypes of human aggression as they relate to childhood psychopathology. METHOD: Critical review of the pertinent literature. RESULTS: In humans, as well as in animals, the term aggression encompasses a variety of behaviors that are heterogeneous for clinical phenomenology and neurobiological features. No simple extrapolation of animal subtypes to humans is possible, mainly because of the impact of complex cultural variables on behavior. On the whole, research into subtypes of human aggression has been rather limited. A significant part of it has been conducted in children. Clinical observation, experimental paradigms in the laboratory, and cluster/factor-analytic statistics have all been used in an attempt to subdivide aggression. A consistent dichotomy can be identified between an impulsive-reactive-hostile-affective subtype and a controlled-proactive-instrumental-predatory subtype. Although good internal consistency and partial descriptive validity have been shown, these constructs still need full external validation, especially regarding their predicting power of comorbidity, treatment response, and long-term prognosis. CONCLUSIONS: Our understanding and treatment of children and adolescents with aggressive behavior can benefit from research on subtypes of aggression. The differentiation between the impulsive-affective and controlled-predatory subtype as qualitatively different forms of aggressive behavior has emerged as the most promising construct. Specific therapeutic hypotheses could be tested in this context and contribute to a full validation of these concepts.

Adolescent↗

A Danish diabetes risk score for targeted screening: the Inter99 study.

OBJECTIVE: To develop a simple self-administered questionnaire identifying individuals with undiagnosed diabetes with a sensitivity of 75% and minimizing the high-risk group needing subsequent testing. RESEARCH DESIGN AND METHODS: A population-based sample (Inter99 study) of 6,784 individuals aged 30-60 years completed a questionnaire on diabetes-related symptoms and risk factors. The participants underwent an oral glucose tolerance test. The risk score was derived from the first half and validated on the second half of the study population. External validation was performed based on the Danish Anglo-Danish-Dutch Study of Intensive Treatment in People with Screen Detected Diabetes in Primary Care (ADDITION) pilot study. The risk score was developed by stepwise backward multiple logistic regression. RESULTS: The final risk score included age, sex, BMI, known hypertension, physical activity at leisure time, and family history of diabetes, items independently and significantly (P<0.05) associated with the presence of previously undiagnosed diabetes. The area under the receiver operating curve was 0.804 (95% CI 0.765-0.838) for the first half of the Inter99 population, 0.761 (0.720-0.803) for the second half of the Inter99 population, and 0.803 (0.721-0.876) for the ADDITION pilot study. The sensitivity, specificity, and percentage that needed subsequent testing were 76, 72, and 29%, respectively. The false-negative individuals in the risk score had a lower absolute risk of ischemic heart disease compared with the true-positive individuals (11.3 vs. 20.4%; P<0.0001). CONCLUSIONS: We developed a questionnaire to be used in a stepwise screening strategy for type 2 diabetes, decreasing the numbers of subsequent tests and thereby possibly minimizing the economical and personal costs of the screening strategy.

Adult↗

Systematic review of the prevention of delayed ischemic neurological deficits with hypertension, hypervolemia, and hemodilution therapy following subarachnoid hemorrhage.

OBJECT: There is uncertainty about the efficacy of hypertension, hypervolemia, and hemodilution (triple-H) therapy in reducing the occurrence of delayed ischemic neurological deficits (DINDs) and death after subarachnoid hemorrhage. The authors therefore conducted a systematic review to evaluate the efficacy of triple-H prevention in decreasing the rate of clinical vasospasm, DINDs, and death. METHODS: The authors systematically reviewed studies identified based on a MEDLINE, EMBASE, and COCHRANE Register search of articles published between 1966 and 2001, and reference lists of identified articles. An independent assessment of each study's methodological quality, population, intervention, and outcomes (rates of symptomatic vasospasm, DINDs, and death) was performed. Summary relative risk estimates were calculated for the main outcomes using fixed- or random-effect models, as appropriate. Only four prospective, comparative studies with a total of 488 patients were identified. The median internal validity score was 0.5 (range 0-2); the median external validity score was 3 (range 2-6). Compared with no prevention, triple-H therapy was associated with a reduced risk of symptomatic vasospasm (relative risk [RR] 0.45, 95% confidence interval [CI] 0.32-0.65), but not DIND (RR 0.54, 95% CI 0.2-1.49). The risk of death was higher (RR 0.68, 95% CI 0.53-0.87). Sensitivity analyses including only randomized, controlled trials showed no evidence of statistically significant results for these major end points. CONCLUSIONS: The paucity of information and important limitations in the design of the studies analyzed preclude evaluation of the efficacy of triple-H prevention and formulation of any recommendations regarding its use for the prevention of cerebral vasospasm.

Blood Pressure↗

Organizational interventions: facing the limits of the natural science paradigm.

This paper reviews current challenges in the conceptualization, design, and evaluation of organizational interventions to improve occupational health. It argues that attempts to confirm cause-and-effect relationships and allow prediction (maximize internal validity) are often made at the expense of generalizability (external validity). The current, dominant experimental paradigm in the occupational health research establishment, with its emphasis on identifying causal connections, focuses attention on outcome at the expense of process. Interventions should be examined in terms of (i) conceptualization, design and implementation (macroprocesses) and (ii) the theoretical mediating mechanisms involved (microprocesses). These processes are likely to be more generalizable than outcomes. Their examination may require the use of both qualitative and quantitative methodologies. It is suggested that such an approach holds unexplored promise for the healthier design, management, and organization of future work.

Occupational Health↗

[Objective evaluation of the cerebral organic-psychiatric axis syndrome--experiences with the encephalopathy questionnaire].

This supplements the publication on "A standardized questionnaire for behavior typical of encephalopathy" (Meyer-Probst, 1978) which dealth with content validation and provisional calibration. The paper is concerned with external validation and, under such headings as "Axis syndrome and milieu factors", "Axis syndrome and school performance" and "Axis syndrome and age", presents the results from questionnaires of an extreme group comparison to illustrate the clinical and diagnostic value of the axis syndrome concept acc. to Göllnitz. Relations of dependence and the usefulness of assessment procedures are discussed, as well as the concept of the axis syndrome in connection with newer trends in specialized psychopathology research.

Brain Damage, Chronic↗

Bias and causal associations in observational research.

Readers of medical literature need to consider two types of validity, internal and external. Internal validity means that the study measured what it set out to; external validity is the ability to generalise from the study to the reader's patients. With respect to internal validity, selection bias, information bias, and confounding are present to some degree in all observational research. Selection bias stems from an absence of comparability between groups being studied. Information bias results from incorrect determination of exposure, outcome, or both. The effect of information bias depends on its type. If information is gathered differently for one group than for another, bias results. By contrast, non-differential misclassification tends to obscure real differences. Confounding is a mixing or blurring of effects: a researcher attempts to relate an exposure to an outcome but actually measures the effect of a third factor (the confounding variable). Confounding can be controlled in several ways: restriction, matching, stratification, and more sophisticated multivariate techniques. If a reader cannot explain away study results on the basis of selection, information, or confounding bias, then chance might be another explanation. Chance should be examined last, however, since these biases can account for highly significant, though bogus results. Differentiation between spurious, indirect, and causal associations can be difficult. Criteria such as temporal sequence, strength and consistency of an association, and evidence of a dose-response effect lend support to a causal link.

Bias↗

General practitioners' decision making for mental health problems: outcomes and ecological validity.

A major problem in the field of medical decision making is the ecological (external) validity of the results. In a Judgement Analysis study on mental health, vignettes were used to capture the decision strategies of 28 General Practitioners (GPs). Two different decision strategies for mental health problems could be distinguished. Although the results were statistically satisfactory and met the assumptions of Judgement Analysis, it was considered necessary to determine the ecological validity of the vignettes. Video tapes (n = 90) of GP consultations were scored in terms of the units of information (cues) which had been used in the vignette study. Additional data gave access to the judgements of the GPs, which were comparable to the judgements obtained in the vignette study. Results showed that the weights given to the different cues in the vignette study were situated within the confidence interval of the weights from the video study. Thus indicating that the results obtained from the vignette study have ecological validity.

Cues↗

Validation of a model predicting spontaneous pregnancy among subfertile untreated couples.

OBJECTIVE: To provide external validation of the Eimers model, which predicts spontaneous pregnancy among subfertile couples within the first year after the definitive establishment of the diagnostic category. DESIGN: Live birth rates predicted by an adapted version of the Eimers model were tested against observed live birth rates in a Canadian cohort study. SETTING: Fertility clinics in university medical centers. PATIENT(S): One thousand sixty-one couples consulting for subfertility due to cervical hostility, male subfertility, or unexplained subfertility. INTERVENTION(S): None. MAIN OUTCOME MEASURE(S): We measured the discriminative ability and reliability of the predictions from the model. RESULT(S): The live birth rate was lower in the Canadian population than in the Eimers population. Overall, the prognostic effect of the predictors did not differ significantly in both populations. The model showed moderate predictive power in the Canadian population. With adjustment of the average live birth rate, the reliability of the model was satisfactory. CONCLUSION(S): The Eimers model gave reliable spontaneous pregnancy predictions in the Canadian validation population after adjustment of the average live birth rate.

Adult↗