Search PubMedSearch

SEARCH · Search PubMed

Results for “Validation Studies as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,791 records · Page 15Linked to original sources

Patient-reported outcome measures for depression or anxiety symptoms in patients with cardiovascular disease: A COSMIN systematic review.

BACKGROUND: Depression and anxiety are common in patients with cardiovascular disease (CVD), but the measurement quality of patient-reported outcome measures (PROMs) used in this population remains unclear. This review aimed to evaluate the methodological quality, measurement properties, and certainty of evidence for depression and anxiety PROMs in adults with CVD and to inform instrument selection. METHODS: Following COSMIN and PRISMA guidance, four databases were searched from inception to February 2026. Studies assessing measurement properties of PROMs in adults with CVD were included. Methodological quality was evaluated using the COSMIN Risk of Bias checklist, and certainty of evidence was graded using an adapted GRADE approach. RESULTS: Sixty-six studies assessing 38 PROMs were included, comprising 29 generic and 9 CVD-specific instruments. Six PROMs met COSMIN Category A criteria: Cardiac Depression Scale-Short Form, Patient Health Questionnaire-9, Beck Depression Inventory-II, Hospital Anxiety and Depression Scale, Generalized Anxiety Disorder-7, and Major Depression Inventory. Four instruments were classified as Category C because of insufficient structural validity. Content-validity evidence was largely indeterminate or of limited certainty. Only 24 studies used confirmatory factor analysis or Rasch analysis, and no study assessed measurement error or responsiveness. Cross-cultural validity evidence was scarce. CONCLUSIONS: Six PROMs met Category A criteria, but selection should remain purpose- and context-specific. Particular attention should be given to somatic symptom overlap and intended clinical use. Further validation should prioritize content validity, measurement invariance, responsiveness, measurement error, and clinimetric performance.

Humans

Tranexamic acid protects human dermal fibroblasts from D-galactose-induced senescence via the GPR30/MAPK pathway.

BACKGROUND: Tranexamic acid (TXA) is widely used for pigmentary disorders, but its anti-ageing potential remains unclear. This study aimed to evaluate whether topical 3% TXA improves early periorbital wrinkles in women with facial melasma and to investigate whether TXA protects human dermal fibroblasts from D-galactose-induced senescence via the GPR30/MAPK pathway. METHODS: Fifty women with melasma were randomized to 3% TXA serum plus moisturizer or moisturizer alone for 8 weeks, with follow-up to week 12. Periorbital wrinkles were graded using a modified Fitzpatrick Wrinkle Scale (MFWS). Separately, D-gal-induced senescence in HDFs was assessed via viability, SA-β-gal activity, senescence markers, ROS, antioxidant enzymes, SASP/ECM gene expression, and MAPK activation. GPR30 involvement was examined using antagonist G15, shRNA knockdown, and molecular docking. RESULTS: Topical TXA produced significantly greater MFWS reductions versus moisturizer alone at weeks 4, 8, and 12, with benefit persisting post-treatment. In HDFs, TXA preserved viability, reduced SA-β-gal positivity, attenuated p21/p16, restored Lamin B1, decreased ROS, and rescued antioxidant activities. TXA downregulated IL-6, IL-8, MMP1, and MMP3, and suppressed D-gal-induced ERK, JNK, and p38 phosphorylation. These effects were weakened by G15 or GPR30 knockdown; docking supported a stable TXA-GPR30 interaction. CONCLUSIONS: TXA showed clinical anti-wrinkle activity in melasma patients and protected HDFs from D-gal-induced senescence, partly via GPR30-dependent modulation of oxidative stress, SASP/ECM expression, and MAPK signalling. TXA is a promising candidate for skin ageing intervention.

Humans

Psychometric Evaluation of the Breast Inflammatory Symptom Severity Index Versions 2 and 3 Among Lactating Women.

OBJECTIVE: To evaluate the psychometric properties of Versions 2 and 3 of the Breast Inflammatory Symptom Severity Index (BISSI). DESIGN: Secondary data analysis of clinical trial data. SETTING: Private physiotherapy practices, a public tertiary hospital, and a community in Melbourne, Australia. PARTICIPANTS: Women more than 7 days after birth with inflammatory conditions of the lactating breast (N = 43). METHODS: We performed confirmatory factor analysis of the BISSI Version 2 to examine item loading, which informed development of the BISSI Version 3 (V3). We assessed convergent validity by comparing total BISSI V3 scores with human milk sodium to potassium ratio (Na+:K+) at Trial Days 1, 3, and 10 using Bland-Altman plots. We compared item-level scores for size of affected area with objective receiver operating characteristic curve analysis to assess discriminant validity for symptom severity and Cronbach's alpha for internal reliability. RESULTS: After confirmatory factor analysis, we removed two items, resulting in a six-item BISSI V3. All retained items demonstrated comparable loading on the overall scale. Limits of agreement for total BISSI V3 scores and item-level scores for size of affected area were acceptable at all time points, with more than 90% of observations falling within 2 standard deviations of the mean difference, supporting convergent validity. Discriminant validity of the BISSI V3 was supported. We found high internal reliability at both time points CONCLUSION: Our findings provide evidence for the validity and reliability of the BISSI V3 and support its continued development and for clinical use of the BISSI V3 and human milk Na+:K+ analysis to enhance management of inflammatory conditions of the lactating breast.

breastfeeding

Holmium laser enucleation of the prostate for the treatment of lower urinary tract symptoms in men with benign prostatic hyperplasia.

RATIONALE: A range of surgical options is available for the treatment of benign prostatic hyperplasia (BPH), including holmium laser enucleation of the prostate (HoLEP). The evidence is unclear regarding differences in functional, perioperative, and morbidity outcomes between these modalities. OBJECTIVES: To assess the effects of holmium laser enucleation of the prostate compared with other surgical treatments for lower urinary tract symptoms in men with benign prostatic hyperplasia. SEARCH METHODS: We searched multiple databases (including MEDLINE, Embase, CENTRAL, Web of Science, LILACS, and the International HTA database), trial registries, and conference abstracts through April 08, 2026. ELIGIBILITY CRITERIA: We only included randomized trials of men over 40 years of age with a prostate volume of at least 20 mL (assessed by digital rectal examination, ultrasound, or conventional imaging) who exhibited lower urinary tract symptoms (LUTS) defined by an International Prostate Symptom Score (IPSS) of eight or greater undergoing surgical interventions for BPH. OUTCOMES: The critical outcomes measured were the urologic symptoms score, the quality-of-life score, and major adverse events. The important outcomes measured were: re-treatment, erectile function, ejaculatory function, transfusions, acute urinary retention, indwelling urinary catheter duration, and hospital stay duration. RISK OF BIAS: We used the Cochrane risk of bias tool (RoB 1) to assess for potential sources of bias on a study and outcome level basis. SYNTHESIS METHODS: We pooled outcome data using the random-effects model and performed meta-analyses using the Mantel-Haenszel method. We assessed statistical heterogeneity in the pooled data by visually inspecting forest plots and using the I2 statistic to quantify it. We used the GRADE framework to assess the certainty of evidence. INCLUDED STUDIES: We included 52 trials that included 6242 participants that compared HoLEP to other surgical interventions for benign prostatic hyperplasia. The median age of participants across the studies ranged from 65 to 74 years. The baseline prostate volume ranged from 30 cc to 142 cc. Baseline IPSS scores ranged from 19.6 to 28.6 (range 0-35). SYNTHESIS OF RESULTS: We prioritized comparing HoLEP with transurethral resection of the prostate (TURP) at short-term follow-up (up to 12 months), because TURP is the long-standing reference standard and the predominant comparator in randomized surgical trials. Findings for the four remaining comparisons (laser ablation, alternative energy source enucleation, other minimally invasive therapies, and simple prostatectomy), for long-term follow-up, and for all remaining outcomes are reported in full in the review. Compared to TURP, at short-term follow-up: Critical outcomes - HoLEP may result in little to no difference in short-term urologic symptom scores measured using the IPSS (range 0 to 35; lower values reflect fewer symptoms) (MD -0.67, 95% CI -1.20 to -0.14; I² = 93%; 14 studies, 1666 participants, low-certainty evidence). - HoLEP may result in little to no difference in short-term quality of life (range 0 to 6; lower values reflect better quality of life) (MD -0.04, 95% CI -0.23 to 0.15; I² = 73%; 6 studies, 876 participants, low-certainty evidence). - HoLEP may result in little to no difference in short-term major adverse events (RR 0.75, 95% CI 0.35 to 1.58; I² = 0%; 10 studies, 1147 participants, low-certainty evidence). Important outcomes - HoLEP likely results in little to no difference in re-treatment (RR 0.45, 95% CI 0.14 to 1.50; I² = 0%; 8 studies, 813 participants, moderate-certainty evidence). - HoLEP likely results in little to no difference in erectile function (MD -0.03, 95% CI -0.47 to 0.42; I² = 0%; 3 studies, 518 participants, moderate-certainty evidence). - Ejaculatory function: we did not find any data for this outcome. - HoLEP likely reduces the need for blood transfusion (RR 0.19, 95% CI 0.09 to 0.42; I² = 0%; 15 studies, 1755 participants, moderate-certainty evidence). AUTHORS' CONCLUSIONS: Compared with TURP, HoLEP may achieve similar relief of urologic symptoms, similar quality of life, and similar rates of major adverse events in the first 12 months after surgery, and probably similar re-treatment rates and erectile function. HoLEP likely reduces the need for blood transfusion; this is the only advantage of HoLEP that the randomized evidence, as summarized here, supports as clinically important. There was insufficient evidence to assess outcomes in the subset of individuals with larger prostates or on anticoagulation. Future research should prioritize long-term trials reporting sexual function and urinary incontinence outcomes, recruit men with very large prostates (≥ 150 cc) or on anticoagulation therapy, and evaluate cost-effectiveness and training requirements. FUNDING: No external funding was received for this review. REGISTRATION: The protocol for this review was published in the Cochrane Database 2019 (https://doi.org/10.1002/14651858.CD013291).

Humans

The Soundtrack of Everyday Life: Real-world Music Listening Habits of Adult Cochlear Implant Users.

OBJECTIVE: Characterize real-world patterns of music listening and reward sensitivity among adult cochlear implant (CI) users compared with normal-hearing (NH) listeners. STUDY DESIGN: Cross-sectional observational study. SETTING: Online. PATIENTS: Adults (&#x2265;18&#xa0;y) with a CI or NH who used a music-streaming platform as their primary listening method. INTERVENTIONS: None. MAIN OUTCOME MEASURES: Objective measures included platform-derived audio features (acousticness, danceability, energy, tempo, and valence), listening volume, unique-song ratio, and decade preferences. Self-reported measures included listening habits and the Barcelona Music Reward Questionnaire (BMRQ). Group comparisons used ANCOVAs and mixed-effect models adjusting for age and gender; within-CI analyses compared prelingual versus postlingual deafness. RESULTS: Among 16 CI users (69% male, 39.0&#xb1;15.6&#xa0;y) and 29 NH listeners (38% male, 32.2&#xb1;11.3&#xa0;y), CI users demonstrated a higher unique-song ratio (&#x3b2;=-0.157, CI as reference; 95%CI [-0.281, -0.034]; P =0.014) and stronger preference for older music (Pillai trace=0.537; F8,35 =5.07; P <0.001), adjusting for age. No significant group differences were observed in weekly listening time, listening volume, audio features, or BMRQ scores (total and subscores). Equivalence was confirmed for Emotion Evocation and Sensory-Motor subscales. There were no statistically significant differences in listening environments after Holm correction. No statistically detectable differences were observed between pre- and postlingually deafened CI users. CONCLUSIONS: Musically active CI users showed no significant differences in listening volume or overall music-reward sensitivity compared with NH peers, but demonstrated higher unique-song ratios and a bias toward older music. Findings highlight the value of ecologically valid data in understanding real-world music experiences among CI users.

Adult

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

A Dynamic Nomogram to Predict Metabolic Dysfunction-Associated Fatty Liver Disease in Patients with Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) involves multiple metabolic disorders. This study aimed to identify high-risk populations for metabolic dysfunction-associated fatty liver disease (MAFLD) in patients with MetS and to establish a dynamic predictive nomogram. METHODS: A total of 627 patients with MetS from six regions in Zhejiang Province were enrolled and categorized into MAFLD and non-MAFLD groups, then randomly assigned to training and validation sets at a ratio of 7:3. Independent predictors of MAFLD were identified using least absolute shrinkage and selection operator regression and multivariable logistic regression analyses. These predictors were then used to construct a dynamic nomogram. RESULTS: A total of 627 patients with MetS were included in the final analysis, of whom 77.0% (483/627) were diagnosed with MAFLD. Multivariable logistic regression analysis identified body mass index (BMI), waist circumference (WC), total cholesterol (TC), alanine aminotransferase (ALT), MetS-defined dysglycemia, and education level as independent risk factors for MAFLD. MetS-defined dysglycemia showed the highest odds ratio (OR) for MAFLD development [OR = 1.87, 95% confidence interval (CI): 1.07-3.29]. Although the number of MetS components and the metabolic syndrome score were significantly associated with MAFLD in univariate analysis, they were not independently associated with MAFLD in the multivariate model. A dynamic nomogram for predicting MAFLD risk in patients with MetS was developed and internally validated. The area under the receiver operating characteristic curve was 0.834 (95% CI: 0.787-0.880) in the training set and 0.839 (95% CI: 0.771-0.899) in the validation set, indicating strong predictive performance. Bootstrap internal validation demonstrated good agreement between predicted and observed outcomes in calibration curves. Decision curve analysis further indicated favorable clinical applicability of the nomogram. CONCLUSION: BMI, WC, TC, ALT, MetS-defined dysglycemia, and education level are independent risk factors for MAFLD. A dynamic nomogram for predicting MAFLD risk in patients with MetS was successfully developed and validated.

Humans

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Meniscal preservation in the age of biologics: toward a quantitative decision algorithm for personalized repair.

BACKGROUND: Despite advances in arthroscopic repair and biologic augmentation, surgical indication for meniscal tears remains heterogeneous. No standardized framework currently integrates biomechanical, clinical, and biological determinants to guide repair versus resection. PURPOSE: To develop a quantitative decision model-the Meniscal Preservation Score (MPS)-that unifies biomechanical and biological evidence to stratify reparability potential and standardize treatment selection in meniscal surgery. METHODS: A systematic evidence synthesis conducted in accordance with PRISMA 2020 reporting standards of studies published from 2000 to 2025 in PubMed, Embase, and Scopus identified key determinants of meniscal healing. Five consistent predictors-patient age, vascularity, tear morphology, associated pathology, and activity profile-were weighted through a two-round modified Delphi consensus among ten experienced knee surgeons. The resulting 0-9-point MPS was incorporated into a stepwise decision tree linking lesion morphology, biological context, and surgical strategy. Conceptual validation used 50 simulated cases and a retrospective cohort of 45 patients to test agreement between algorithm recommendations and expert surgical decisions. RESULTS: The MPS achieved 86% concordance with expert judgment in simulation and 84% agreement in clinical validation. In this retrospective exploratory cohort, cases in which surgical management was concordant with MPS recommendations demonstrated higher mean IKDC scores at 24&#xa0;months and lower observed reoperation rates. These findings should be interpreted as associative rather than causal, as treatment allocation was not controlled and discordant cases may have represented inherently more complex pathology. CONCLUSION: The MPS represents an evidence-informed decision-support framework designed to systematize reparability assessment. While exploratory analyses suggest structural coherence with expert reasoning, prospective implementation and external validation are required before clinical adoption as a predictive tool. LEVEL OF EVIDENCE: conceptual model with exploratory validation.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Urinary Small Extracellular Vesicle DNA as a Biomarker for the Non-Invasive Diagnosis of Bladder Cancer.

Existing diagnostic technologies for bladder cancer (BC) suffer from low sensitivity, low specificity, or a lack of validation. Therefore, validated, non-invasive diagnostic biomarkers with high sensitivity and specificity for early detection of BC are needed to complement and improve upon the limitations of existing diagnostic methods. We used low-pass whole genome sequencing (LP-WGS) technology to detect copy number variations (CNVs) in small extracellular vesicle (sEV) DNA isolated from urine samples of patients. Based on these results, we constructed and validated a diagnostic model to differentiate between benign and malignant bladder lesions. We conducted a receiver operating characteristic analysis and calculated the area under the curve (AUC) to evaluate the performance of the diagnostic model. The urine sEV-DNA LP-WGS data revealed CNV differences between benign and malignant samples. The diagnostic model achieved an AUC of 0.953, a sensitivity of 86.7%, and a specificity of 100% in the training cohort and an AUC of 0.985, a sensitivity of 90%, and a specificity of 100% in the validation cohort. Even at the lowest coverage depth of 0.01X, the performance of the diagnostic model remained relatively robust. Notably, the performance of this diagnostic model surpassed that of the biomarker neuron-specific enolase (sensitivity: 85.7% vs. 64.3%; specificity: 100% vs. 87.5%) and urinary cytology (sensitivity: 100% vs. 66.7%; specificity: 100% vs. 94.1%). Our study demonstrates that urine sEV-DNA exhibits high discriminatory power in distinguishing between benign and malignant bladder lesions, making it a promising tool for auxiliary diagnosis of BC.

Humans

Long-term hormone therapy for perimenopausal and postmenopausal women.

BACKGROUND: Hormone therapy is widely provided to control menopausal symptoms and has been used for the management and prevention of cardiovascular disease, osteoporosis and dementia in older women. This is an updated version of a Cochrane review first published in 2005. OBJECTIVES: To assess the long-term effects of prolonged use (at least one year) of hormone therapy on mortality, cardiovascular outcomes, cancer, gallbladder disease, fractures and cognition in perimenopausal and postmenopausal women. SEARCH METHODS: We used the Cochrane Gynaecology and Fertility Group Specialised Register, CENTRAL, MEDLINE, three other databases and two trial registers, together with reference checking, citation searching and contact with study authors to identify the studies included in the review. The latest search date was 26 September 2024. SELECTION CRITERIA: We included randomised, double-blind trials in which peri- or postmenopausal women took hormone therapy or placebo for at least one year. We included various oestrogen formulations, with or without progestogens. We focused on studies assessing hormone therapy's effects on long-term clinical outcomes, including death, coronary events and cancer. Hormone therapy's efficacy in managing menopausal symptoms was beyond the scope of this review, and is assessed in other Cochrane reviews. DATA COLLECTION AND ANALYSIS: Two review authors independently selected studies, assessed risk of bias and extracted data. We calculated risk ratios (RRs) for dichotomous data and mean differences (MDs) for continuous data, along with 95% confidence intervals (CIs). We assessed the certainty of the evidence using GRADE. MAIN RESULTS: We included 24 studies - with two newly added in this update - involving 45,660 participants. We derived nearly 70% of the data from two well-conducted studies: the Heart and Estrogen/progestin Replacement Study (HERS 1998) and the large, multi-component Women's Health Initiative research programme, which included two hormone therapy arms (WHI 1998). Across all the studies, most participants were postmenopausal American women with one or more comorbidities. The mean participant age in most studies was over 60 years. Only one included study focused on perimenopausal women. We present full results for all included studies with available data in the main review. The results presented below are drawn from WHI 1998, in which the combined hormone therapy arm and the oestrogen-only arm were run concurrently, with women assigned to the appropriate trial based on their uterus status. One study with 16,608 postmenopausal women with an intact uterus compared combined continuous hormone therapy (conjugated equine oestrogen and medroxyprogesterone acetate) to placebo, and measured outcomes at an average of 5.6 years of follow-up. Based on this study, combined continuous hormone therapy probably makes little to no difference to the risk of a coronary event (RR 1.17, 95% CI 0.95 to 1.44; moderate-certainty evidence). It may increase the risk of stroke (RR 1.39, 95% CI 1.09 to 2.09; low-certainty evidence) and venous thromboembolism (RR 2.03, 95% CI 1.55 to 6.64; low-certainty evidence). Compared to placebo, combined continuous hormone therapy probably increases the risk of breast cancer (RR 1.27, 95% CI 1.03 to 1.56; moderate-certainty evidence) and probably makes little to no difference to the risk of lung cancer (RR 1.06, 95% CI 0.77 to 1.46; moderate-certainty evidence). It may increase gallbladder disease requiring surgery (RR 1.64, 95% CI 1.30 to 2.06; 14,203 participants; low-certainty evidence), and probably reduces the risk of all clinical fractures (RR 0.78, 95% CI 0.71 to 0.86; moderate-certainty evidence). One study including 10,739 postmenopausal women who had undergone a hysterectomy compared oestrogen-only (conjugated equine oestrogen) hormone therapy to placebo, and measured outcomes at an average of seven years' follow-up. Based on this study, oestrogen-only hormone therapy probably makes little to no difference to the risk of coronary events (RR 0.94, 95% CI 0.78 to 1.13), venous thromboembolism (RR 1.32, 95% CI 1.00 to 1.74) and breast cancer (RR 0.79, 95% CI 0.61 to 1.01), all with moderate-certainty evidence. It may make little to no difference to the risk of lung cancer (RR 1.04, 95% CI 0.73 to 1.48; low-certainty evidence). Oestrogen-only hormone therapy probably increases the risk of stroke (RR 1.33, 95% CI 1.06 to 1.67) and gallbladder disease requiring surgery (RR 1.78, 95% CI 1.42 to 2.24), and probably reduces the risk of all clinical fractures (RR 0.73, 95% CI 0.65 to 0.80), all with moderate-certainty evidence. We judged most included studies to have a low risk of bias for most domains. The overall certainty of evidence for the main comparisons was moderate. The main limitation was that only about 30% of women were 50 to 59 years old at baseline, the age group most likely to consider hormone therapy for vasomotor symptoms. AUTHORS' CONCLUSIONS: Long-term follow-up of women using hormone therapy suggests that the risk profiles vary between combined hormone therapy and oestrogen-only therapy. Oestrogen-only hormone therapy probably makes little to no difference to coronary events, and probably increases the risk of stroke and gallbladder disease. It probably makes little to no difference in the risk of breast cancer, and probably reduces the risk of all fractures. Combined hormone therapy may increase the risk of thromboembolism and probably increases the risk of breast cancer. These results should be interpreted with caution as they are based on one study using oral hormone therapy, which may not represent the risks of the hormone therapy currently used in clinical practice.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

A pragmatic randomized controlled trial of self-directed online writing interventions for posttraumatic stress symptoms in a real-world digital setting.

Background: Public health and other large-scale crises, such as the COVID-19 pandemic, have intensified the global mental health burden, creating unprecedented demand for accessible interventions for posttraumatic stress symptoms (PTSS).Objective: We evaluated the feasibility and effectiveness of two self-directed online writing interventions embedded within China's WeChat ecosystem during the COVID-19 pandemic through a pragmatic randomised controlled trial.Methods: Between December 2021 and August 2022, 1,526 adults were screened for PTSS via a Tencent Medinfo Mini-Program. Eligible participants (n&#x2009;=&#x2009;211) were randomised to Guided Narrative Technique-Writing (GNT-W, n&#x2009;=&#x2009;100) or Expressive Writing (EW, n&#x2009;=&#x2009;111). Both interventions comprised three self-directed daily writing sessions delivered entirely online without human support. Primary outcome was PTSD symptom severity (PTSD Checklist-Short), assessed at baseline, post-intervention, 2-week, and 1-month follow-ups.Results: While initial engagement followed typical digital health patterns (64.5% overall attrition), participants who initiated treatment showed strong adherence (77% completion). Both interventions were associated with significant within-group reductions in PTSS severity (GNT-W: b&#x2009;=&#x2009;-0.43, p&#x2009;=&#x2009;.023, d&#x2009;=&#x2009;-0.43; EW: b&#x2009;=&#x2009;-0.60, p&#x2009;=&#x2009;.001, d&#x2009;=&#x2009;-0.58), with no significant between-group difference (group &#xd7; time: b&#x2009;=&#x2009;0.18, p&#x2009;=&#x2009;.48). GNT-W did not confer additional benefit over EW protocol on PTSS severity.Conclusions: Both self-directed writing interventions were associated with within-group reductions in PTSS; without an inactive control condition, however, these changes cannot be firmly attributed to the interventions. GNT-W showed no advantage over the simpler EW protocol. These findings offer preliminary support for embedding scalable, low-barrier writing interventions in widely used digital platforms.Chinese Clinical Trial Registry: ChiCTR2000034836.

Humans

Magnesium sulphate for women at risk of preterm birth for neuroprotection of the fetus.

BACKGROUND: Magnesium sulphate is a common therapy in perinatal care. Its benefits when given to women at risk of preterm birth for fetal neuroprotection (prevention of cerebral palsy for children) were shown in a 2009 Cochrane review. Internationally, use of magnesium sulphate for preterm cerebral palsy prevention is now recommended practice. As new randomised controlled trials (RCTs) and longer-term follow-up of prior RCTs have since been conducted, this review updates the previously published version. OBJECTIVES: To assess the effectiveness and safety of magnesium sulphate as a fetal neuroprotective agent when given to women considered to be at risk of preterm birth. SEARCH METHODS: We searched Cochrane Pregnancy and Childbirth's Trials Register, ClinicalTrials.gov, and the World Health Organization (WHO) International Clinical Trials Registry Platform (ICTRP) on 17 March 2023, as well as reference lists of retrieved studies. SELECTION CRITERIA: We included RCTs and cluster-RCTs of women at risk of preterm birth that assessed prenatal magnesium sulphate for fetal neuroprotection compared with placebo or no treatment. All methods of administration (intravenous, intramuscular, and oral) were eligible. We did not include studies where magnesium sulphate was used with the primary aim of preterm labour tocolysis, or the prevention and/or treatment of eclampsia. DATA COLLECTION AND ANALYSIS: Two review authors independently assessed RCTs for inclusion, extracted data, and assessed risk of bias and trustworthiness. Dichotomous data were presented as summary risk ratios (RR) with 95% confidence intervals (CI), and continuous data were presented as mean differences with 95% CI. We assessed the certainty of the evidence using the GRADE approach. MAIN RESULTS: We included six RCTs (5917 women and their 6759 fetuses alive at randomisation). All RCTs were conducted in high-income countries. The RCTs compared magnesium sulphate with placebo in women at risk of preterm birth at less than 34 weeks' gestation; however, treatment regimens and inclusion/exclusion criteria varied. Though the RCTs were at an overall low risk of bias, the certainty of evidence ranged from high to very low, due to concerns regarding study limitations, imprecision, and inconsistency. Primary outcomes for infants/children: Up to two years' corrected age, magnesium sulphate compared with placebo reduced cerebral palsy (RR 0.71, 95% CI 0.57 to 0.89; 6 RCTs, 6107 children; number needed to treat for additional beneficial outcome (NNTB) 60, 95% CI 41 to 158) and death or cerebral palsy (RR 0.87, 95% CI 0.77 to 0.98; 6 RCTs, 6481 children; NNTB 56, 95% CI 32 to 363) (both high-certainty evidence). Magnesium sulphate probably resulted in little to no difference in death (fetal, neonatal, or later) (RR 0.96, 95% CI 0.82 to 1.13; 6 RCTs, 6759 children); major neurodevelopmental disability (RR 1.09, 95% CI 0.83 to 1.44; 1 RCT, 987 children); or death or major neurodevelopmental disability (RR 0.95, 95% CI 0.85 to 1.07; 3 RCTs, 4279 children) (all moderate-certainty evidence). At early school age, magnesium sulphate may have resulted in little to no difference in death (fetal, neonatal, or later) (RR 0.82, 95% CI 0.66 to 1.02; 2 RCTs, 1758 children); cerebral palsy (RR 0.99, 95% CI 0.69 to 1.41; 2 RCTs, 1038 children); death or cerebral palsy (RR 0.90, 95% CI 0.67 to 1.20; 1 RCT, 503 children); and death or major neurodevelopmental disability (RR 0.81, 95% CI 0.59 to 1.12; 1 RCT, 503 children) (all low-certainty evidence). Magnesium sulphate may also have resulted in little to no difference in major neurodevelopmental disability, but the evidence is very uncertain (average RR 0.92, 95% CI 0.53 to 1.62; 2 RCTs, 940 children; very low-certainty evidence). Secondary outcomes for infants/children: Magnesium sulphate probably resulted in little to no difference in severe intraventricular haemorrhage (grade 3 or 4) (RR 0.81, 95% CI 0.64 to 1.04; 6 RCTs, 6542 infants; moderate-certainty evidence) and may have resulted in little to no difference in chronic lung disease/bronchopulmonary dysplasia (average RR 0.92, 95% CI 0.77 to 1.10; 5 RCTs, 6689 infants; low-certainty evidence). Primary outcomes for women: Magnesium sulphate may have resulted in little or no difference in severe maternal outcomes potentially related to treatment (death, cardiac arrest, respiratory arrest) (RR 0.32, 95% CI 0.01 to 7.92; 4 RCTs, 5300 women; low-certainty evidence). However, magnesium sulphate probably increased maternal adverse effects severe enough to stop treatment (average RR 3.21, 95% CI 1.88 to 5.48; 3 RCTs, 4736 women; moderate-certainty evidence). Secondary outcomes for women: Magnesium sulphate probably resulted in little to no difference in caesarean section (RR 0.96, 95% CI 0.91 to 1.02; 5 RCTs, 5861 women) and postpartum haemorrhage (RR 0.94, 95% CI 0.80 to 1.09; 2 RCTs, 2495 women) (both moderate-certainty evidence). Breastfeeding at hospital discharge and women's views of treatment were not reported. AUTHORS' CONCLUSIONS: The currently available evidence indicates that magnesium sulphate for women at risk of preterm birth for neuroprotection of the fetus, compared with placebo, reduces cerebral palsy, and death or cerebral palsy, in children up to two years' corrected age. Magnesium sulphate may result in little to no difference in outcomes in children at school age. While magnesium sulphate may result in little to no difference in severe maternal outcomes (death, cardiac arrest, respiratory arrest), it probably increases maternal adverse effects severe enough to stop treatment. Further research is needed on the longer-term benefits and harms for children, into adolescence and adulthood. Additional studies to determine variation in effects by characteristics of women treated and magnesium sulphate regimens used, along with the generalisability of findings to low- and middle-income countries, should be considered.

Humans

Thymosin-&#x251;1 for people with chronic hepatitis B.

RATIONALE: Chronic hepatitis B is a global public health concern. It is caused by infection with the hepatitis B virus (HBV). The goal of treating chronic HBV infection is to prevent progression to chronic hepatitis, cirrhosis, hepatic decompensation, liver failure, hepatocellular carcinoma, and death. Individual studies have evaluated various immunomodulatory therapies with inconsistent results. Thymosin-&#x251;1 is known to have antiviral effects; however, results of randomised clinical trials on the effects of thymosin-&#x3b1;1 as a potential treatment for people with chronic HBV have been inconsistent. OBJECTIVES: To assess the benefits and harms of thymosin-&#x251;1 therapy in people with chronic hepatitis B. SEARCH METHODS: We searched the Cochrane Hepato-Biliary Group Controlled Trials Register, CENTRAL, MEDLINE, four other databases and six trials registers, in addition to reference checking, citation searching, and contacting study authors to identify trials for inclusion. The latest search date was 10 June 2026. ELIGIBILITY CRITERIA: We included randomised controlled trials (RCTs) that evaluated thymosin-&#x3b1;1 at any dose, route of administration, or formulation type, in people with chronic hepatitis B regardless of age, sex, or ethnicity. Thymosin-&#x3b1;1 could have been administered as monotherapy, in combination with an additional drug, or in addition to standard medical treatment and compared with placebo, no intervention, the same additional drug, or the same standard medical treatment. OUTCOMES: Our critical outcomes were all-cause mortality, serious adverse events, and health-related quality of life. Among our important outcomes were HBV-related morbidity, HBV-related mortality, non-serious adverse events, and the proportion of people without histological improvements. RISK OF BIAS: We used the Cochrane Risk of bias 2 tool (RoB 2) to assess risk of bias. SYNTHESIS METHODS: We followed Cochrane methods. We conducted meta-analyses for predefined outcomes using data from the longest follow-up period, irrespective of the risk of bias judgements. We presented dichotomous outcome results as risk ratios (RRs) and continuous outcome results as mean differences, with 95% confidence intervals (CIs) at their longest follow-ups. We used the random-effects model for our primary analyses. We used GRADE to assess the certainty of the evidence for each outcome. INCLUDED STUDIES: We included 10 RCTs conducted in Bangladesh, China, Italy, Korea, Singapore, and Taiwan, with 1349 randomised participants (range: 12 to 690; 1045 (77.5%) were male). Among the trials reporting age, none included participants younger than 17 years (age range: 17 to 75 years). The trials were published between 1991 and 2018, and assessed thymosin-&#x251;1 in adults with chronic hepatitis B infection, with or without comorbidities. Only two trials mentioned comorbidities (cirrhosis and acute-on-chronic liver failure). The trials compared thymosin-&#x251;1, with or without a cointervention, with placebo or no intervention, or with the same cointervention. The control interventions were placebos in two trials and no intervention in two. The remaining six trials administered co-interventions, such as interferon, pegylated interferon, lamivudine, and standard medical therapy (entecavir or tenofovir), and entecavir. Follow-ups ranged from six months to five years after the end of treatment (median: 12 months). Four trials were funded by industry, five by research grants, and one provided no information. All 10 trials (11 records) provided data on at least one outcome in our review. We identified no ongoing trials. Sixteen studies are awaiting assessment due to incomplete reporting. We received no responses to our enquiries. SYNTHESIS OF RESULTS: Thymosin-&#x251;1, compared with the control interventions, may reduce all-cause mortality (RR 0.53, 95% CI 0.29 to 0.96; I&#xb2; = 0%; 3 studies, 907 participants; very low-certainty evidence), serious adverse events (RR 0.72, 95% CI 0.53 to 0.99; I&#xb2; = 0%; 5 studies, 1056 participants; low-certainty evidence), HBV-related mortality (RR 0.53, 95% CI 0.29 to 0.96; I&#xb2; = 0%; 3 studies, 907 participants; very low-certainty evidence), non-serious adverse events (RR 0.47, 95% CI 0.27 to 0.83; I&#xb2; = 0%; 5 studies, 300 participants; very low-certainty evidence), and may have little to no effect on health-related quality of life (MD 0.70, 95% CI -2.55 to 3.95; I&#xb2; not applicable; 1 study, 161 participants; very low-certainty evidence; score range: 0 to 100; the higher the score, the better) and on histological improvement (RR 0.51, 95% CI 0.13 to 2.06; I&#xb2; = 74%; 2 studies, 702 participants; very low-certainty evidence). The evidence is very uncertain about the effect of thymosin-&#x251;1 on hepatitis B-related morbidity (RR 0.86, 95% CI 0.54 to 1.40; I&#xb2; = 3%; 3 studies, 854 participants; very low-certainty evidence). We judged the certainty of evidence to be low for serious adverse events and very low for the remaining outcomes. Reasons for downgrading were mainly due to study limitations, including overall high or some concerns for risk of bias; imprecision of the pooled effect estimates (including wide or very wide confidence intervals crossing the line of no effect, and small participant numbers); and inconsistency due to substantial heterogeneity (I&#xb2; = 74%). The test for subgroup differences provided no evidence of differences in effect according to thymosin&#x2011;&#x3b1;1 administration for any outcome (P &#x2265; 0.05). AUTHORS' CONCLUSIONS: We assessed the certainty of evidence as very low for all outcomes except for serious adverse events (low). Therefore, we are not sure whether thymosin-&#x3b1;1 monotherapy versus placebo or no intervention, or with the same co-interventions, reduces all-cause mortality, serious adverse events, HBV-related mortality, and non-serious adverse events, nor whether it has any effect on quality of life (based on one trial) and histological improvement. The effect of thymosin-&#x251;1 on HBV-related morbidity is very uncertain. We observed no statistically significant differences between trials with and without cointerventions. We found no ongoing trials. FUNDING: This Cochrane review had no dedicated funding. REGISTRATION: Protocol available via DOI: 10.1002/14651858.CD014610.

Humans