Search PubMedSearch

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis. A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST + AI for prediction model studies. Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST + AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection. AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans

Predictive Models for Hypoglycemia Risk in Haemodialysis Patients With Diabetic Kidney Disease: Systematic Review and Meta-Analysis.

AIM: To provide evidence for selecting and developing reliable clinical assessment tools for hypoglycemia in diabetic kidney disease patients during haemodialysis. DESIGN: Review. METHODS: Systematic searches were performed in 9 Chinese and English databases to collect literature regarding the development of hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease. Two reviewers independently performed literature screening, data extraction, risk-of-bias assessment, and applicability evaluation. The Prediction Model Risk of Bias Assessment Tool was used to assess the risk of bias and applicability of the included studies. Meta-analysis was conducted using R software. DATA SOURCES: CNKI, Wanfang, VIP, CBM, PubMed, Cochrane Library, EMbase, Web of Science, and CINAHL. The search period covered from the establishment date of each database to December 2025. RESULTS: Six studies, comprising six prediction models, were included. Two studies performed internal validation, and three conducted external validation. All models reported the area under the curve, ranging from 0.813 to 0.866, and calibration measures. Four studies were rated as having a high risk of bias, while all six demonstrated good overall applicability. The meta-analysis showed that the pooled AUC value of the six studies was 0.846 (95% CI: 0.823-0.867). CONCLUSION: Research on hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease remains in the developmental stage. Although the included prediction models exhibited satisfactory apparent discriminatory ability and clinical applicability, most of the original studies suffered from a high risk of bias and lacked adequate validation. The true predictive performance and clinical application value of these models remain to be further verified. Accordingly, routine and unconditional clinical application is not recommended at this stage. Future studies should include more high-quality, multicenter external validation and develop models with high generalizability, favourable clinical applicability, and robust predictive performance to facilitate early identification of hypoglycemia risk in this population. IMPACT: This study systematically evaluated the hypoglycemia risk prediction models for diabetic kidney disease patients during haemodialysis, and the research on hypoglycemia risk prediction models for maintenance haemodialysis patients during dialysis is still in the development stage. This study provides a reference for clinical medical staff to select or develop hypoglycemia risk prediction and assessment tools for diabetic kidney disease patients during haemodialysis. REPORTING METHOD: This study was conducted in accordance with the relevant guidelines of the EQUATOR Network and followed the TRIPOD-SRMA Checklist. PATIENT OR PUBLIC CONTRIBUTION: No patient or public contribution. TRIAL REGISTRATION: PROSPERO: CRD420251243352.

Humans

Cross-Platform Proteomics and Machine Learning Algorithms Nominate Plasma Biomarkers of Stroke Diagnosis.

BACKGROUND: Blood-based biomarkers for stroke subtyping could improve triage in emergency settings. We used cross-platform proteomics to identify plasma biomarkers differentiating major stroke diagnostic groups. METHODS: We conducted a case-control study using 2 biorepositories. Plasma was collected in the emergency department from adults with suspected stroke before therapeutic intervention. Differentially enriched proteins were identified across acute ischemic stroke, intracerebral hemorrhage, transient ischemic attack, and stroke mimics using SomaScan discovery proteomics (Grady). Differentially enriched proteins were nominated using pairwise and multigroup comparisons and adjusted for clinical covariates. Protein panels were created using least absolute shrinkage and selection operator logistic regression. Internal validation used repeated nested cross-validation (rCV) and targeted mass spectrometry (MS), while external validation used data-independent acquisition &#xa0;mass spectrometry in an independent cohort (Yale). RESULTS: We included 100 subjects (40 with acute ischemic stroke, 20 with intracerebral hemorrhage, 20 with transient ischemic attack, 20 with stroke mimics) in discovery and 80 subjects (20 per group) in external validation cohorts. SomaScan quantified 7307 proteins, of which 61 differentiated stroke subtypes. We identified 7 protein classifiers for acute ischemic stroke (rCV-area under the curve, 0.82 [95% CI, 0.78-0.86]), 6 for intracerebral hemorrhage (rCV-area under the curve, 0.70 [95% CI, 0.64-0.76]), 8 for transient ischemic attack (rCV-area under the curve, 0.78 [95% CI, 0.73-0.84]), and 7 for stroke mimics (rCV-area under the curve, 0.81 [95% CI, 0.77-0.86]). Targeted proteomics internally validated 11 proteins, and data-independent acquisition-mass spectrometry externally validated 32 proteins, including VTN (vitronectin), PLG (plasminogen), and S100A9 as top stroke mimics, transient ischemic attack, and intracerebral hemorrhage classifiers. CONCLUSIONS: This study highlights plasma proteomics as a valuable tool for discovering protein biomarkers of stroke diagnosis. These findings support further validation in larger, multicenter cohorts to facilitate biomarker-guided stroke diagnosis in acute care.

Humans

The impact of methodological factors on child psychotherapy outcome research: a meta-analysis for researchers.

Two recent meta-analyses have generated evidence for child and adolescent psychotherapy effects. However, critics note that such meta-analyses often include studies with methodological shortcomings which might invalidate their results. In the present study, we explored whether the results of the most extensive child/adolescent meta-analysis might have been influenced by such methodological variables, focusing on internal validity and external validity factors. Together, these factors accounted for two-thirds as much variance as the substantive factors (e.g., type of therapy, age) in the original meta-analysis. This suggests that relative to these therapy and child-characteristic variables, methodological factors have a substantial, though smaller, impact on meta-analysis results. In general, increased experimental rigor was related to larger effect sizes; this argues against the hypothesis that methodologically weak studies have led to an overestimate of therapy effects. No significant interactive relations were found between validity factors and predictors of outcome; this suggests that the relations noted in previous meta-analyses between outcome and various variables were not distorted by the validity factors tested here.

Adaptation, Psychological

Prediction equations for the estimation of body composition in the elderly using anthropometric data.

To study the relationship between health and nutritional status in elderly populations, information about body composition is essential. To collect this information in large epidemiological studies, practical methods based on anthropometric data must be available. In the present study the relationship between body composition, determined by densitometry, and anthropometric data in 204 elderly men and women, aged 60-87 years, was analysed. Existing prediction equations described in the literature, and mainly based on young and middle-aged subjects, generally underestimated percentage body fat in the elderly study population. Therefore, new prediction equations were developed, based on sex and the sum of two (biceps and triceps) or four (biceps, triceps, suprailiaca and subscapula) skinfolds or the body mass index (BMI). Addition of age or body circumferences to the models did not improve the prediction of body density. Internal cross validation and external validation revealed that the formulas are valid for the estimation of body density in elderly subjects. The standard errors of estimate of the three models, expressed as percentage body fat, were 5.6, 5.4 and 4.8% respectively.

Aged

The structure of the Mental Health Inventory among Chinese in Taiwan.

This study attempted to ascertain the construct validity and external validity of the Mental Health Inventory in a Chinese population in Taiwan and contrast these results with results obtained from studies of several U.S. populations. In particular, a series of measurement models were specified and evaluated to address the issues of reliability and validity. Data were collected from personal interviews of a probability sample of 1,194 Chinese respondents 14 years of age and older in four townships in southwest Taiwan. The Mental Health Inventory was found to involve two major components: positive well-being and psychological distress. As a hierarchical structure, each component consists of one second-order and two or three first-order factors. The relationships between well-being and distress can be characterized as substantially independent and modestly bipolar depending on the level and specification.

Adolescent

Development of an interview-based geriatric depression rating scale.

The geriatric depression rating scale (GDRS) is a new interview-based depression rating scale designed for use with adults 60 years of age or older. The scale was developed to fill a need for an instrument that would be sensitive to the problems encountered in assessing depression among older adults. The GDRS was designed by using items from the self-report Geriatric Depression Scale (GDS) as topic areas in a structured clinical interview similar to that of the Hamilton Rating Scale for Depression (HRSD). The 35-item rating scale was administered to 68 older individuals with a range of affective disturbance. The scale was found to have internal consistency and split-half reliability comparable to the HRSD and GDS. Concurrent validity, construct validity, external criterion validity, sensitivity, and specificity were all found to be acceptable.

Aged

A critique of project evaluations.

In recent years an increased stress has been placed on the evaluation of mental health, education, and welfare service programs. The majority of studies readily available to most evaluators represent local project evaluations which usually contain diverse references to different aspects of the evaluative process. For evaluative results to be even minimally useful to other projects, however, certain requirements must be met. These are: (1) internal validity, (2) external validity, (3) specification of the population and treatment being implemented, and (4) standardization of indicators of treatment impact. To determine the extent to which published project impact evaluations meet these criteria, a study was undertaken to "evaluate the evaluations" themselves within heroin addiction treatment programs. Six high-yield journals and 100 random sources were systematically searched for reports of evaluations which provided measures of success in terms of the consumer. Articles were analyzed in regard to our four prerequisites for cross-project comparisons regarding process variables, impact variables, and methodologies. It became clear, however, that our original objectives in evaluating either the usefulness of published project evaluations or testing any specific impact hypotheses were not achievable due to the state of evaluative measurement and reporting practices at this time. The major problems we eoncountered in our inability to complete a necessary and potentially fruitful comparative assessment of project evaluations are discussed in detail with recommendations for future work.

Heroin Dependence

Analysis of end-stage renal disease mediated by cuproptosis-related genes.

OBJECTIVE: The complex pathophysiological mechanism of end-stage renal disease (ESRD) has not been fully understood. Cuproptosis is a newly discovered type of programmed cell death. Therefore, this study attempts to clarify the relationship between cuproptosis-related genes (CRGs) and the phenotype of ESRD. MATERIALS AND METHODS: The National Center for Biological Information Gene Expression Omnibus database was applied to obtain the GSE37171 dataset comprising whole-genome microarray analysis of peripheral blood samples. A 3&#xa0;:&#xa0;1 case-control design was employed with 75 ESRD patients and 20 healthy controls who were frequency-matched for age, sex, and ethnicity. Based on differentially expressed genes (DEGs) and genes related to cuproptosis, CRGs were identified. Thereafter, we explored two different subpopulations based on the cuproptosis gene and analyzed their expression and immune infiltration. Genes specific to the CRG cluster were identified through the weighted gene co-expression network analysis algorithm, and the best prediction model was determined and verified by four machine learning methods. RESULTS: The study identified 14 differentially expressed CRGs, among which ATP7B, SLC31A1, LIAS, LIPT1, DLD, MTF1, CDKN2A, DBT, and DLST had relatively high expression levels in the ESRD samples. Compared with the control group, expression levels of FDX1, DLAT, PDHA1, PDHB, and GLS were significantly lower in the ESRD group, and CRGs played a key role in the regulation of immune infiltration in ESRD. Two cuproptosis-related molecular clusters were identified in the ESRD samples. Cluster2 was more correlated with the immune infiltration of ESRD. By analyzing the intersection points between CRG cluster and key genes of ESRD, a total of 888 specific DEGs were identified. Functional differences related to specific DEGs were further explored using gene set variation analysis. Five significant genes (SMC5, USP47, USP53, AGA, and DMXL1) were identified by the support vector machine model as key predictors for ESRD disease risk, achieving an area under the curve (AUC) of 1.00 in internal validation. However, external validation in independent cohorts is required prior to clinical application. Individual gene analysis showed an AUC >&#xa0;0.81 in discriminating ESRD patients from healthy controls, and the expression of all 5 genes in ESRD patients was significantly lower than in the control group. CONCLUSION: This study clarified the relationship between CRGs and the phenotype of ESRD, analyzed their specific roles in the immune microenvironment, and obtained a predictive model, providing new insights for the study of its potential therapeutic targets.

Humans

A methodological framework for the design of research on the evaluation of residents.

This paper describes a construct validation framework for research on the selection and evaluation of residents. The application of the proposed methodology to surgery residents is described. The need to measure non-cognitive and neuropsychological factors in addition to cognitive knowledge and technical ability is emphasized, and a research strategy that integrates theory formulation, internal validation, and external validation is presented. In this context, residents' competence is viewed as a multivariate construct that requires validation through longitudinal empirical studies and the use of multivariate statistical approaches.

Clinical Competence

Scientific challenges in the application of randomized trials.

In recent years, scientific challenges in the application of randomized trials have become more apparent, especially with the extension of such trials to the assessment of nondrug treatments, such as health education, psychotherapy, and health care provision. Six issues (individual v group randomization, blinding and unblinding, the effect of trial participation on outcome, selective subject participation, treatment compliance, and standardized v individualized treatment) are discussed in terms of their impact on internal validity, generalizability (external validity), and clinical relevance. Specific design strategies may be necessary to enhance these methodological and clinical desiderata. Attention to these challenges should lead to improvements in future randomized trials.

Attitude

Could the preoperative urethral curve be used to predict immediate urinary continence following Retzius-sparing robot-assisted radical prostatectomy? A retrospective multi-center study.

PURPOSE: Immediate urinary continence (UC) recovery following Retzius-sparing robot-assisted radical prostatectomy (RS-RARP) remains highly variable, highlighting the need for reliable preoperative prediction. We aimed to develop and validate models to identify patients likely to achieve immediate UC recovery following RS-RARP. MATERIALS AND METHODS: A total of 580 prostate cancer patients who underwent RS-RARP from four medical centers were assigned to a training set (n=348), an internal validation set (n=103) and an external validation set (n=129). Independent predictors were identified through univariate analysis and LASSO regression. A nomogram was constructed using multivariate logistic regression. Its performance was evaluated with receiver operating characteristic (ROC) curve, calibration curves, and decision curve analysis. RESULTS: Immediate UC recovery was observed in 84.5% (294/348) of patients in the training cohort, 80.6% (83/103) in the internal validation cohort, and 81.4% (105/129) in the external validation cohort, respectively. Multivariate analysis identified membranous urethral length (MUL) (OR=1.23, P=0.029) and urethral curvature (OR=2.84, P<0.001) as independent predictors, while prostate volume (PV) (OR=0.84, P <0.001) as a protective factor. The nomogram integrating MUL, PV, and urethral curvature demonstrated superior predictive accuracy, with an AUC of 0.87 (95% CI, 0.83-0.91) in the training cohort. The bootstrap-corrected calibration slope was 0.96, and the Brier score was 0.08.&#xa0;Calibration curves and decision curve analysis confirmed the predictive accuracy and clinical utility of the nomogram. CONCLUSIONS: Our study introduces a novel quantitative method for assessing urethral curvature. The mpMRI-based model, integrating urethral curvature and prostate spatial configuration, offers enhanced predictive accuracy for postoperative immediate UC recovery.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Prediction of incident heart failure in established atherosclerotic cardiovascular disease: the SMART2-HF model.

BACKGROUND AND AIMS: Patients with established atherosclerotic cardiovascular disease (ASCVD) are at high risk of developing heart failure (HF). However, incident HF is not part of the risk assessment of current guideline-recommended models. The aim of this study was to develop and externally validate the SMART2-HF model for prediction of incident HF in patients with ASCVD. METHODS: SMART2-HF was developed in 7698 individuals with established ASCVD (coronary, cerebrovascular, or peripheral artery disease, or abdominal aortic aneurysm) but without prior HF from the UCC-SMART cohort. Cox proportional hazards models including sex-predictor interactions and with age as the time scale were derived to estimate the 10-year and lifetime risk of incident HF (hospitalization for HF or HF-related death), accounting for competing non-HF mortality. Predictors, limited to routinely available clinical characteristics, were aligned with the SMART2 risk model for recurrent cardiovascular (CV) risk in the same population. External validation was performed in 240 741 patients with ASCVD from six data sources: the Clinical Practice Research Datalink, the HUNT3 study, the SWEDEHEART Registry, the ASCVD-Particles cohort, the Estonian Biobank and the international REACH Registry. RESULTS: During a median follow-up of 11.2 years (interquartile range 6.1-16.4 years), 1031 incident HF events (13%) occurred in the UCC-SMART cohort. In the external validation data sources, a total of 24 885 incident HF events (10%) occurred. The pooled C-statistic was .696 (95% confidence interval .674-.717), with consistent performance in subgroups by sex and type of ASCVD. Predicted risks matched observed incidence in external validation. CONCLUSIONS: The SMART2-HF model enables the prediction of incident HF in patients with ASCVD. Aligned with the guideline-recommended SMART2 model for recurrent CV risk, SMART2-HF can be used as a complementary tool in this population.

Humans

Is it clinically possible to distinguish nonhemorrhagic infarct from hemorrhagic stroke?

BACKGROUND AND PURPOSE: Diagnosis of the nonhemorrhagic ischemic type of stroke by analysis of patients' clinical features is considered unreliable because no clinical feature is specific. The diagnosis is so difficult to establish that we cannot hope to use the same method to make a reliable diagnosis in all stroke cases. In this study, we propose a simple scoring system with a positive predictive value of close to 100% to distinguish nonhemorrhagic infarct from hemorrhagic stroke. This scoring is available for all physicians in bedside diagnosis even if this score can be applied to a subgroup of patients. METHODS: Twenty-six clinical variables that might potentially distinguish cerebral hemorrhage from infarction were recorded in patients consecutively admitted to our stroke unit for stroke lasting more than 24 hours with at least unilateral motor weakness affecting face and/or arm and/or leg (internal validity study). Patients previously receiving anticoagulant therapy were excluded. We used CT scan as the gold standard. We used multivariate logistic regression to establish a clinical score from which we derived the classification rule. This rule was validated with data from the next 200 consecutive patients hospitalized in the stroke unit (external validity study). RESULTS: Three hundred sixty-eight patients were enrolled in the internal study. The obtained score was (2 x alcohol consumption) + (1.5 x plantar response) + (3 x headache) + (3 x history of hypertension)--(5 x history of transient neurological deficit)--(2 x peripheral arterial disease)--(1.5 x history of hyperlipidemia)--(2.5 x atrial fibrillation on admission). All patients with a score less than 1 (n = 123) had a nonhemorrhagic infarct (ie, 40% of the 305 patients with a nonhemorrhagic infarct). No threshold was found to diagnose cerebral hemorrhage with a sufficiently high positive predictive value. Among the 200 patients enrolled in the external validity study, 72 patients with a score below 1 had a nonhemorrhagic infarct (ie, 43% of patients with a nonhemorrhagic infarct). CONCLUSIONS: Diagnosis of nonhemorrhagic infarct can be made in 36% (95% confidence interval [CI], 29 to 43) of patients with a high level of accuracy (100% in the external validity study, which gives a 95% CI of 93 to 100). Thus, 43% (95% CI, 36 to 50) of patients with a nonhemorrhagic infarct could receive a bedside diagnosis. The score is simple and can be calculated from information available to all physicians.

Adult

[Quasi experimental evaluation of public health interventions (author's transl)].

The classic experiment, the randomised controlled trial, is the best known and most revered of evaluation research methods. Randomization in community-based intervention trials, however, is not always possible because of ethical problems arising from with holding the experimental treatment from the control groups or the difficulties in conducting experiments in field settings which do not approach controlled laboratory conditions. In such circumstances, quasi-experimental or observational designs must be used. Two major principles are involved in using quasi-experimental methods: (1) the logic for establishing causality between treatment and effect is the same as that for randomised experiments, but the problems of assessing causality or internal validity are greater, and (2) assessment of the external validity or generalizability of quasi-experimental findings crucial to the interpretation of results. Selected quasi-experimental designs using time series and comparison groups are described with examples from public health intervention trials where threats to internal validity have been assessed by using different analytic techniques or gathering additional evidence. Quasi-experimental evaluations are most useful when opportunities exist for testing rival hypotheses concerning the internal and external validity, or the findings can be used to complement true experiments.

Epidemiologic Methods

Risk prediction in patients with heart failure with preserved ejection fraction: the LIFE-Preserved model.

BACKGROUND AND AIMS: Heart failure (HF) with preserved ejection fraction (HFpEF) constitutes a heterogeneous disease with varying prognosis. Given the rising incidence of HFpEF, accurate risk prediction for these patients is needed to identify high-risk individuals, who may benefit the most from preventive treatments. The LIFE-Preserved model was developed and validated for the prediction of individual short-term and lifetime risk for HF hospitalization or cardiovascular (CV) death in patients with HFpEF. METHODS: LIFE-Preserved was derived in 20 332 patients aged 40-90 years with a left ventricular ejection fraction &#x2265; 50% from the Swedish HF Registry. Cause- and sex-specific Cox models were derived to predict the risk of HF hospitalization or CV death using 14 routinely available predictors. Use of age as the timescale allowed for predictions beyond the maximum follow-up duration in the derivation data, adjusted for competing risks. External validation was performed in two trials (EMPEROR-Preserved and TOPCAT-Americas) and three registries (NHS England Secure Data Environment, Veterans Affairs, and HF-Particles). Model performance was assessed by discrimination and calibration. RESULTS: During a median follow-up of 1.8 years (interquartile range .6-4.2, maximum 19 years), 9341 first HF hospitalizations or CV deaths (46%) were observed in Swedish HF Registry. External validation included data from 28 062 patients with HFpEF [9930 (35%) first HF hospitalizations or CV deaths]. Pooled C-statistics were .714 (95% confidence interval .652-.775) in trials and .658 (95% confidence interval .599-.717 in registries, with adequate calibration in all external validation sources. Performance was similar in men and women. An interactive calculator of the LIFE-Preserved model has been made available here. CONCLUSIONS: The LIFE-Preserved model enables prediction of short-term and lifetime risk of HF hospitalization or CV death in patients with HFpEF. The model could serve as a tool to identify high-risk HFpEF patients, guiding clinical management and shared decision-making.

Humans