Search PubMedSearch

SEARCH · Search PubMed

Results for “validation studies”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,498 recordsLinked to original sources

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75 161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et al., Nanda et al., Naylor et al., and Van Leeuwen et al., each showing fair discrimination. The Teede et al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et al. and van Leeuwen et al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Transcranial Photobiomodulation Variables Assessment Battery: Development and Validation.

Transcranial photobiomodulation (tPBM) response variability is partly driven by biophysical characteristics such as skin tone and hair properties that attenuate photon penetration, and by lifestyle factors including sleep quality, alcohol use, and nicotine consumption that disrupt the mitochondrial and vascular pathways on which tPBM acts. To date, no validated self-report tool exists to capture these moderators systematically. To address this gap, the tPBM Variables Assessment Battery was developed and psychometrically evaluated. It integrates adapted versions of established measures (Brief Pittsburgh Sleep Quality Index, E-cigarette Dependence Scale, Hair Scale Assessment PRO, Monk Skin Tone Scale, and Heaviness of Smoking Index), validated wellbeing evaluators (Ryff's Psychological Wellbeing), and custom measures (Hairstyle Classification, Hair Color Classification). Face and content validity met recommended expert thresholds, internal consistency was acceptable across adapted subscales, and criterion validity analyses confirmed meaningful associations between the lifestyle components and PROMIS-10 global health outcomes. The battery is low-burden, digitally deployable, and psychometrically defensible, offering a practical tool for characterizing the variables most likely to moderate tPBM response in home-use studies.

Humans

Assessing the Concurrent Validity of the Australian Treatment Outcomes Profile in a Methamphetamine Dependent Treatment-Seeking Population.

INTRODUCTION: The Australian Treatment Outcomes Profile (ATOP) is a brief clinical tool assessing substance use, health and well-being used in Australian alcohol and other drug treatment services. It is validated for use with clients using alcohol, opioids and cannabis, but not yet for clients who primarily use methamphetamine. METHODS: An embedded validation study was undertaken in treatment-seeking adults enrolled in a randomised double-blind placebo-controlled trial of lisdexamfetamine for methamphetamine dependence with sites in New South Wales, South Australia and Victoria. Participant demographics were collected during study screening. The ATOP and comparators (Time Line Follow Back, Opiate Treatment Index, Depression Anxiety Stress Scale, WHOQOL-BREF and Personal Wellbeing Index) were collected at baseline. Continuous ATOP items were analysed using Pearson's correlation coefficient, and dichotomous items were analysed using Fleiss's &#x3ba;. Agreement was rated as strong where measures were &#x2265;&#x2009;0.50, moderate where agreement was 0.30-0.49, and weak where <&#x2009;0.30. RESULTS: One hundred and eighteen study participants (2018-2020) had data for concurrent validity analysis. Strong validity was demonstrated for physical health, psychological health, quality of life, injecting drug use and crime items, and for days of use for amphetamines, alcohol, cannabis and cocaine. There was weak validity for days of use for benzodiazepines. Heroin use days and other opioid use days were endorsed by fewer than five participants and were therefore unable to be assessed. DISCUSSION AND CONCLUSIONS: The ATOP is valid for use in a treatment-seeking methamphetamine-dependent population, expanding the range of tools for assessment and standardised outcome monitoring across different settings and services.

Humans

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC&#x2009;=&#x2009;0.91-0.92 [0.83-0.97]; k&#x2009;=&#x2009;55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC&#x2009;=&#x2009;0.90-0.91 [0.85-0.95] (k&#x2009;=&#x2009;228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs&#x2009;=&#x2009;0.90 [0.83-0.94] (k&#x2009;=&#x2009;124) and ICC&#x2009;=&#x2009;0.91 [0.72-0.98] (k&#x2009;=&#x2009;9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load&#x2013;velocity relationship

Predictive Validity of Violence Screening Tools in Emergency and Psychiatric Services: A Systematic Review.

Violence against healthcare staff, including a threat or an act of violence toward people during their work, poses a physical and psychological risk to workers internationally. Screening is an important strategy in preventing violence against healthcare professionals. The aim of this systematic review was to synthesize evidence on the predictive validity of risk assessment tools used to screen for violence and aggression risk toward healthcare workers in emergency and psychiatric departments (PD). Primary studies that examined the predictive validity of risk assessment tools for workplace violence were identified via a systematic search of Medline, PsycINFO, Embase, and the Cochrane databases. There were 62 eligible studies, ten of which had a lower risk of bias (RoB). Those studies with high RoB were primarily due to a failure to present calibration measures as part of the analysis. All included studies adopted a longitudinal design and were conducted in PDs. The ten highest-quality studies reported on eight different instruments, four of which showed acceptable to outstanding predictive performance. The Dynamic Appraisal of Situational Aggression and the Br&#xf8;set Violence Checklist showed the best predictive performance; they were also validated in emergency departments and are best suited for short-term risk prediction. We recommend that the selection of a risk assessment tool should consider the following: (a) the target population, (b) the violence operationalization, and (c) the purpose of the monitoring. We note that the use of a screening tool should be a part of a multicomponent strategy to ensure staff safety.

Humans

Direct background subtraction LC-MS/MS assay for human plasma progesterone: Full validation and comparative application.

OBJECTIVE: To develop and validate a liquid chromatography-tandem mass spectrometry method based on direct background subtraction for the quantification of endogenous progesterone in human plasma. METHODS: Protein precipitation was used for sample preparation with deuterated progesterone as the internal standard. Chromatographic separation was performed on an ACQUITY C18 column using gradient elution with 0.1% formic acid in water and acetonitrile at a flow rate of 0.3&#xa0;mL/min. Mass spectrometry was operated in positive electrospray ionization mode with multiple reaction monitoring. Instead of using analyte-stripped matrix or surrogate matrix, authentic plasma was directly used for all validation experiments. Quantitation was achieved by subtracting the background signal, and results were compared with those from the classical method using stripped matrix. RESULTS: Excellent linearity was achieved over 0.1-100&#xa0;ng/mL (R2&#xa0;&#x2265;&#xa0;0.99). Precision, accuracy, recovery, matrix effect, and stability all met FDA and ICH M10 acceptance criteria. Compared with the classical method, the bias in Cmax and AUC0-t was within &#xb1;15%, indicating no significant difference between the two methods. CONCLUSION: The direct background subtraction method avoids laborious preparation of blank matrix, eliminates matrix effect discrepancies, and is simple, efficient, and low-cost. It can serve as a general strategy for endogenous substance determination.

Humans

Low-burden metrics for monitoring healthy diets among nonpregnant females aged 15 to 49 years: a multicountry validation analysis using quantitative 24-hour dietary intake data.

BACKGROUND: Limited nationally representative quantitative dietary intake data and a lack of consensus on lower-burden tools and metrics hinder high-frequency monitoring of healthy diets globally. OBJECTIVES: This study aimed to evaluate the comparative construct validity and potential complementarity of low-burden metrics of a healthy diet among nonpregnant females aged 15 to 49 y. METHODS: Quantitative 24-h dietary intake data collected from 77,118 adolescent and adult females across 27 countries were used to construct low-burden metrics and reference metrics of dietary intake. Associations between mean-standardized low-burden measures or indicators and reference metrics were assessed using linear and logistic mixed-effect models, with Spearman's &#x3c1; used for survey-level rank correlations. Test characteristics identified low-burden indicators best differentiated adherence to reference indicators. RESULTS: An indicator reflecting nonconsumption of sweet foods and/or sweet beverages was most robustly associated with greater adherence to <10% energy from free sugars in upper-middle-income countries {odds ratio [OR] [95% confidence interval (CI)]: 5.35 [5.05, 5.66]}. Food group diversity score (FGDS) was most strongly associated with and differentiated higher mean adequacy ratio of micronutrients [&#x3b2; of 1-standard deviation (SD) change: &#x223c;11 percentage points (9, 12); &#x3c1;: 0.79], whereas noncommunicable disease-Protect score best reflected consumption of &#x2265;400 g/d of fruits and vegetables [range OR of 1-SD changes (95% CI): 2.56-3.01 (2.40, 3.13) in lower-middle and high-income countries, respectively; &#x3c1;: 0.56]. FGDS and Global Diet Quality Score Positive were most consistently associated with achieving &#x2265;25 g/d of fiber and &#x2265;3510 mg/d of potassium across contexts. CONCLUSIONS: Low-burden data collection tools yield valid metrics, enabling high-frequency monitoring of healthy diets across contexts. Specifically, avoiding sweet foods and/or sweet beverages is an indicator for adherence to WHO free sugar guidelines among nonpregnant females in upper-middle-income countries, whereas metrics reflecting nutritious food group diversity strongly reflect better micronutrient adequacy and adherence to WHO guidelines for fruits and vegetables, fiber, and potassium intakes within and across contexts.

Humans

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation

Validation of a Turkish Translation of the Stress in Emergency Healthcare Professionals: The Stress Factors and Manifestations Scale.

AIM: The primary duties of emergency healthcare professionals (EHPs) are to provide emergency patient care to acutely ill and injured individuals. Due to the nature of their work, EHPs operate under constant stress, often requiring rapid decision-making, swift action, and the delivery of necessary medical care in life-or-death situations, sometimes under inadequately safe conditions. Therefore, the aim of this study is to determine the validity and reliability of the Emergency Healthcare Professional Stress Factors and Symptoms (SEHP:SFMS) Scale in Turkish for identifying stress factors and symptoms in emergency medical care professionals providing emergency patient care services. DESIGN: A methodological study design was used in this study. METHODS: The study was conducted with the participation of 211 EHPs from employees working in emergency care institutions affiliated with the Mu&#x11f;la Provincial Health Directorate between November 2023 and June 2024. Data were collected via a face-to-face survey. Data were analysed using Lawshe content validity ratio, Kaiser-Meyer-Olkin coefficient, Bartlett test, exploratory factor analysis, principal component analysis, Varimax factor rotation method, confirmatory factor analysis, Cronbach's &#x3b1; internal consistency coefficient, convergent validity, discriminant validity, test-retest, and Spearman correlation coefficient tests. RESULTS: The linguistic translation and cultural adaptation of the SEHP:SFMS showed strong performance. The scope validity index of the scale is 0.83. The item-total correlation values of the scale were found to be between 0.486 and 0.794, and the factor loadings were between 0.474 and 0.816. Confirmatory factor analysis fit indices: &#x3c7;2&#x2009;=&#x2009;248.727; df&#x2009;=&#x2009;101; n&#x2009;=&#x2009;211; p&#x2009;=&#x2009;0.000; &#x3c7;2/df&#x2009;=&#x2009;2.463; RMSEA&#x2009;=&#x2009;0.083; CFI&#x2009;=&#x2009;0.914, SRMR&#x2009;=&#x2009;0.052, which was found to be compatible and acceptable with the proposed 3-factor model. The Cronbach's &#x3b1; reliability coefficient of the scale was 0.931, and the total variance was 61.97%. CONCLUSIONS: SEHP:SFMS is a valid and reliable tool to assess stress factors and symptoms of Turkish emergency healthcare professionals. Its use improves the quality of emergency care. PATIENT OR PUBLIC CONTRIBUTION: These study findings have been used to create a tool with Turkish validity and reliability that allows for the examination of stress factors among healthcare professionals working in emergency and critical services. Identifying and reducing stress factors among healthcare professionals is crucial for the delivery of quality healthcare services. It can also be used to develop targeted interventions and ongoing strategies to facilitate improved clinical supervision and mentoring. IMPLICATION FOR NURSING PRACTICE: Nurses in emergency departments, which are among the most stressful, dynamic, intense, life-saving, and critical environments in healthcare institutions, and where life-saving treatment is administered, are at high risk of experiencing psychological trauma. Trauma experienced in the work environment is a significant problem for nursing. The consequences of trauma negatively affect nurses and institutions. Studies show that post-traumatic stress, anxiety, depression, and burnout are commonly observed in emergency department nurses. In this sense, understanding the stress and stress factors experienced by nurses can guide future interventions. The results of this study are considered important in making visible the stress and stress factors experienced by nurses in the emergency department, and also in guiding managers and nurses working in this field in terms of preventive and protective measures.

Humans

Metformin Adherence and Risk of Polyneuropathy in Type 2 Diabetes Mellitus: An International Matched Cohort Study with Independent Validation.

BACKGROUND: Metformin is a popular first-line glucose-lowering medication for type 2 diabetes mellitus (T2DM). Although metformin reduces the risks of various complications of diabetes, its potential to cause polyneuropathy by depleting vitamin B12 levels is concerning. This study investigated whether the adherence or discontinuation of metformin after adding-on a second-line antiglycemic agent increases the risk of polyneuropathy in patients with T2DM. METHODS: Data from TriNetX were obtained, and patients with T2DM who were receiving second-line antiglycemic agents were divided into metformin-adherent and metformin-nonadherent groups based on prescription claims data. Neuropathy incidence was evaluated using diagnostic claims and nerve conduction examinations. For independent confirmation and external validation of the primary findings, we used data from the National Health Insurance Research Database (NHIRD) of Taiwan. RESULTS: After matching, 58,027 patients were included in each group. Compared with metformin adherent patients, metformin nonadherent patients had a higher risk of polyneuropathy (adjusted hazard ratios [aHR] 1.26; 95% confidence interval [CI] 1.23-1.29; P < 0.001). Risks of diabetic foot ulcer, amputation, neuropathy-related medication use, and bone fracture were also higher among nonadherent patients. Sensitivity analyses confirmed the robustness of findings. In the validation NHIRD cohort (31,384 matched pairs), metformin nonadherence remained associated with increased polyneuropathy risk (aHR 1.25; 95% CI 1.10-1.42; P < 0.001). CONCLUSIONS: Metformin adherence in patients with T2DM who require second-line treatment may reduce the risk of polyneuropathy; vitamin B supplementation may enhance this benefit.

Humans

Methods for defining equity-stratifying variables: a systematic review of validation studies.

BACKGROUND AND OBJECTIVE: Disease burden is often disproportionally higher among those who are socially disadvantaged by factors defined in the PROGRESS-Plus framework (ie, Place of residence, Race/ethnicity/culture/language, Occupation, Gender/sex, Religion, Education, Socioeconomic status, and Social capital, with "Plus" covering features like age and disability). The accuracy and applicability of case definitions to identify these variables from administrative and clinical health data are unknown. We conducted a systematic review to explore how equity-stratifying variables, as categorized by the PROGRESS-Plus framework, have been defined and validated in epidemiologic studies using administrative health, population-level, or electronic health record (EHR) data. METHODS: Medline, EMBASE, CINAHL, Web of Science, and Google Scholar were searched from the inception of the databases to 2024 for validation studies of equity-stratifying variables in adults using administrative health datasets, health registries, or EHR data. Titles and abstracts, followed by relevant full-text articles, were screened in duplicate by two reviewers for eligibility. The data sources utilized, algorithms employed, and their associated performance measures were extracted and synthesized from included studies. Given substantial heterogeneity in study design, equity-stratifying variable definition, and performance metrics, meta-analysis was not possible. RESULTS: Of the 9099 unique citations screened, 188 full texts were reviewed and 116 were included in this review. Most studies were published between 2019 and 2024 (n = 64, 55%) and were validation studies of race/ethnicity definitions that used race/ethnicity codes or surname list algorithms (n = 66, 57%). No studies examined religion. Regarding the reported performance measure estimates, the race/ethnicity/culture/language equity-stratifying variables category had the largest variability across sensitivity, positive predictive value (PPV), and Cohen's Kappa. Occupation validation studies had the lowest variation in sensitivity and PPV. CONCLUSION: Despite an increasing number of publications reporting on the validation of equity-stratifying variables relevant to the PROGRESS-Plus framework, performance measures varied widely across studies. The significant heterogeneity in equity-stratifying variable definitions and methods used to validate them support the need for further rigorous validation of equity-stratifying variables in administrative and clinical health data. PLAIN LANGUAGE SUMMARY: Disease burden is often higher in people who experience financial hardships, lower level of education, discrimination due to race/ethnicity, and unstable housing. These social factors can be considered health equity factors and are important for understanding health inequalities. Health researchers often use large datasets, such as hospital or electronic health records (EHRs), to study these health equity factors. However, it is not clear how accurately these data sources capture information about people's social circumstances and how these factors are defined. In this study, we reviewed existing research to understand how health equity factors have been defined across health data sources and how accurate they are at measuring aspects of health equity and social disadvantage. Of the more than 9000 studies we identified, we included 116 that met our criteria for this systematic review. Most included studies focused on identifying race and ethnicity, often using codes or surname-based methods. We found that the accuracy of these methods varied widely across studies, meaning results may not always be reliable or comparable. Overall, our findings show that there are inconsistencies in how social factors are defined and measured in health data. This makes it difficult to fully understand and address health inequalities using routinely collected health data. More work is needed to develop and validate better quality and more consistent methods for capturing these important social factors.

Humans

Teaching Engagement and Caregiving Help in the Intensive Care Unit (TEACH-ICU) Scale: Content Validity.

BACKGROUND: Having family members provide care to their loved ones in the intensive care unit (ICU) is a beneficial yet seldom implemented approach. For family members to perform caregiving, nurses must be willing to teach, and such willingness is a developing area of research. OBJECTIVES: To adapt an instrument validated in family members, the Family Willingness for Caregiving Scale, to address nurses' willingness to teach family members caregiving skills. METHODS: Purposive and snowball sampling were used to recruit 10 expert ICU nurses through the American Association of Critical-Care Nurses' research website and social media platforms. The researchers conducted cognitive interviews with the nurses to address the instrument's content validity. RESULTS: The scale was refined based on the participants' feedback. Items were deleted, added, and revised. Furthermore, scale instructions were adjusted to emphasize the willingness to teach families of patients receiving mechanical ventilation. Qualitative themes emerged related to barriers to family engagement, including time constraints, patient acuity, and nurse and family characteristics. CONCLUSIONS: Content validity of the scale was assessed, with future research aimed at pilot testing and evaluating construct validity before using the scale as a research instrument. Practical implications include using the scale as an evaluation tool to determine nurses' willingness to teach family members about caregiving. After evaluation, various strategies could be incorporated to enhance family engagement in adult ICUs.

Humans

Development and Validation of a Predictive Model for Identification of Cognitive Impairment Risk in Older Adults with Subjective Cognitive Decline&#xff1a;A Longitudinal Study.

BACKGROUND: Subjective cognitive decline (SCD) is a transitional state between objective cognitive impairment and cognitively intact mental status, providing a critical window for implementing preventive interventions to delay objective cognitive decline. AIMS: We aimed to develop a predictive model for SCD progression in older adults with mild cognitive impairment (MCI). This model will facilitate the identification of risk factors and establishment of targeted interventions for community-based SCD management. METHODS: Data from the China Health and Retirement Longitudinal Study (CHARLS) was utilized in this study, extracting 18 indicators. Potential predictors selected through univariate Cox regression and LASSO regression analyses were sequentially incorporated into a multivariable Cox regression model. A nomogram was constructed to establish a predictive model. Model validation encompassed Area Under Curve (AUC) metrics for discriminative capacity, complemented by quantitative assessments using calibration curve analysis for precision verification and decision curve analysis (DCA) for clinical utility evaluation. RESULTS: A total of 1099 older adults with SCD were included in the final analysis, of whom 114 (10.3%) developed MCI. Multivariable Cox regression identified residence, marital status, educational level, social participation, gait speed, and baseline cognitive function. The model demonstrated time-dependent AUC values of 0.885, 0.830, 0.839, and 0.836 in the training set when evaluating discriminative capacity at 2-, 4-, 7-, and 9-year, respectively. The predictive model showed excellent predictive ability according to AUC, calibration curve, and DCA. CONCLUSIONS: A predictive model was created to estimate the risk of developing MCI in older individuals with SCD, offering clinician-actionable intervention benchmarks for preventive care.

Humans

An individualized nomogram for predicting progression-free survival in systemic anaplastic large cell lymphoma: a multicenter, retrospective, and internally validated study.

OBJECTIVES: To develop an individualized nomogram for predicting disease progression risk in systemic anaplastic large cell lymphoma (sALCL). METHODS: Independent predictors of progression-free survival (PFS) were identified using Cox regression in a multicenter retrospective cohort of 109 sALCL patients (2010-2022). These were incorporated into a three-factor nomogram, evaluated via bootstrapped internal validation (1000 resamples), ROC analysis, C-index, decision curve analysis (DCA), and clinical impact curve (CIC). RESULTS: A total of 29 PFS events occurred during a median follow-up of 31 months. Multivariable modelling selected serum &#x3b2;2-microglobulin elevation, extranodal disease, and front-line chemotherapy choice (CHOP versus CHOPE or BV+CHP) as autonomous progression drivers. Upon internal bootstrap validation, the nomogram yielded strong prognostic accuracy, achieving AUCs of 0.81, 0.85 and 0.87 for 1-, 3- and 5-year progression-free survival, alongside a corrected C-index of 0.779 (95% CI: 0.699 - 0.861). Calibration plots showed close agreement between predicted and observed outcomes, while DCA confirmed superior net clinical benefit versus conventional IPI or Ann Arbor stratification across multiple decision thresholds. CONCLUSION: This first sALCL-specific nomogram integrates clinical and treatment variables to provide personalized PFS risk estimation. While internally validated, this exploratory, observation-based tool requires external validation and recalibration in prospective cohorts before clinical implementation.

Humans

Predictive evolutionary genomics: principles, validation, and practice.

Climate change and habitat loss are driving rapid evolutionary responses in populations world-wide, which creates an urgent need for evolutionary forecasting in conservation and agriculture. Such forecasting can be categorized into three time scales: trait-based models that use multivariate quantitative genetic equations to project correlated phenotypic responses up to c.&#xa0;20 generations, allele-based analyses that model allele frequency dynamics up to 100 generations, and composite adaptation scores that aggregate many small effects to yield predictions across longer horizons. However, these approaches have remained largely disconnected. Here, we present a Bayesian framework that integrates these three complementary approaches for evolutionary prediction. Our framework combines genomic, phenotypic, and environmental data to yield probabilistic predictions with explicit uncertainty. We show how predictive evolutionary forecasts can be validated with experimental evolution, field experimentation, historical specimens, and reciprocal transplants. These validated forecasts can help advance conservation and agricultural programmes by helping predict which populations are at risk of future extinction, optimizing breeding programmes for future climates, and planning ecosystem management under environmental change. By supporting a shift towards more predictive approaches in evolutionary biology, this framework may help improve our ability to manage biodiversity and food security in a changing world.

Genomics

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

Validation of the lung immune prognostic index in extensive-stage small cell lung cancer: Post hoc analysis of the caspian and IMpower133 phase 3 trials.

BACKGROUND: The Lung Immune Prognostic Index (LIPI) is an inflammation-based biomarker associated with outcomes to immunotherapy across several tumor types. Its prognostic value in extensive-stage small-cell lung cancer (ES-SCLC), however, remains insufficiently validated. We aimed to validate the prognostic impact of LIPI in ES-SCLC using data from two phase III trials. METHODS: Patients enrolled in the CASPIAN (NCT03043872) and IMpower133 (NCT02763579) trials were included. LIPI groups were defined as good (dNLR<3 and LDH<ULN), intermediate (dNLR&#x2265;3 or LDH&#x2265;ULN) and poor (dNLR&#x2265;3 and LDH&#x2265;ULN). Overall survival (OS) and progression-free survival (PFS) were assessed across LIPI categories and treatment arms. RESULTS: LIPI was available for 1140 patients (Good: 34%, Intermediate: 49%, Poor: 17%), including 708 treated with chemotherapy-immunotherapy and 432 with chemotherapy alone. Poor LIPI was associated with unfavorable characteristics, including lower albumin levels and higher rate of liver metastases. Median OS was 14.6 months (95%CI: 12.4-15.9) for LIPI Good, 10.9 (10.1-11.5) for Intermediate, and 8.4 (7.1-9.3) for Poor (p&#x202f;<&#x202f;0.0001). In multivariate models adjusted on gender, age, ECOG, treatment arm and metastatic sites, LIPI remained an independent prognostic factor for OS (HR Poor vs. Good: 1.76, 95%CI: 1.45-2.15, p&#x202f;<&#x202f;0.001) and PFS (HR: 1.59, 95%CI: 1.33-1.90, p&#x202f;<&#x202f;0.001). Although patients with poor LIPI derived limited benefit from immunotherapy, no significant treatment-LIPI interaction was observed. CONCLUSION: This large post hoc analysis confirms LIPI as a robust and clinically applicable prognostic biomarker in ES-SCLC. Patients with poor LIPI have substantially worse outcomes and limited benefit from immunotherapy, highlighting the need for novel therapeutic strategies in this subgroup.

Humans

Self-Report Health Screening Tools in Female Athletes: A Systematic Review of Domain Coverage, Validation, and Use Across Participation Levels.

BACKGROUND: Female athlete health encompasses multiple interconnected domains; however, the self-report screening tools used to assess these domains have not been comprehensively synthesised. OBJECTIVE: To systematically identify self-report health screening tools used to assess female athlete health, map domain coverage, determine validation reporting, and describe application across participation levels. METHODS: This systematic review was pre-registered with PROSPERO ( CRD420251056910 ) and conducted in accordance with PRISMA guidelines. Four databases (PubMed, MEDLINE, SPORTDiscus and Web of Science) were searched from inception to January 2026 using female health and screening-related terms. Methodological quality was appraised using Joanna Briggs Institute and National Institutes of Health tools, and findings were synthesised descriptively. Eligible, peer-reviewed studies reported the use, development or validation of self-report health screening tools assessing one or more domains relevant to female health applied in female athlete populations, spanning recreational through elite participation levels. All sports and activities were included. The search was restricted to English language with no date limits. RESULTS: In total, 360 studies (1990-2026) representing 134,506 female participants spanning recreational to elite sport and 273 screening tools were included. Mental health (n&#x2009;=&#x2009;77, 34.1%), disordered eating (n&#x2009;=&#x2009;33, 14.6%) and body image (n&#x2009;=&#x2009;30, 13.3%) predominated. Domains related to female health, including menstrual health, pelvic floor health, pregnancy/postpartum and breast health were comparatively underrepresented. Most&#xa0;studies reported tools were used for risk identification (n&#x2009;=&#x2009;323,&#xa0;80.3%). Validation reporting was inconsistent, with half (n&#x2009;=&#x2009;180,&#xa0;50%) reporting use of at least one validated tool. Tool use was concentrated in professional and elite sport, with limited inclusion of recreational, masters and disability athlete cohorts. Health literacy constructs were explicitly&#xa0;assessed in 12.5% of studies&#xa0;(n&#x2009;=&#x2009;45). CONCLUSIONS: Health screening in female athlete populations remains fragmented and uneven in domain coverage, with inconsistent validation reporting. Development of integrated, multi-domain and contextually inclusive screening frameworks is warranted.

Journal Article