Search PubMedSearch

SEARCH · Search PubMed

Results for “validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

340 records · Page 3Linked to original sources

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Vitamin B12 Deficiency in Sickle Cell Disease: Method-Driven Estimates and Systematic Diagnostic Misclassification.

OBJECTIVES: To determine whether the reported 0%-70% prevalence of vitamin B12 deficiency in sickle cell disease (SCD) reflects true population variation or diagnostic misclassification. METHODS: We conducted a PRISMA 2020-compliant systematic review of observational studies (January 1, 2000-May 13, 2026; PROSPERO CRD420251087800) assessing B12 status in SCD. PubMed, AJOL, and Google Scholar were searched with citation tracking and dual screening. Diagnostic validity was assessed across biomarker strategy, analytical platform, thresholds, and confounder control using a proposed context-integrated framework to classify methodological robustness and discordance. RESULTS: Fourteen studies were included (57% high-income; 43% LMIC). The evidence base was dominated by limited diagnostic approaches: 71% used immunoassays, over one-third relied on circulating B12 alone, and functional biomarkers were inconsistently applied without systematic confounder adjustment. Prevalence estimates were strongly influenced by diagnostic methods rather than underlying population biology, ranging from 0% to 70% in single-marker studies (mostly 0%-7.1%, with outliers ~50%-70%) and 6.9%-53% in multi-marker studies. Discordance was substantial and greater in LMIC settings than HIC. CONCLUSION: Current diagnostic approaches in SCD appear method-dependent, generating heterogeneous prevalence estimates with uncertain clinical validity. These findings challenge existing estimates and have implications for clinical practice, research design, and diagnostic equity. TRIAL REGISTRATION: ClinicalTrials.gov identifier: CRD420251087800.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Specific Instruments for Caregiving Competence Among Family Caregivers of Cancer Patients: A COSMIN Systematic Review of Psychometric Properties.

OBJECTIVE: To evaluate and summarize the psychometric properties of specific instruments for caregiving competence among family caregivers of cancer patients. METHODS: Systematically searched eight databases for studies published up to November 2025. The methodological quality and psychometric properties of the instruments were evaluated using COSMIN 2.0. Evidence grades were rated using the modified GRADE system (four grades: "High," "Moderate," "Low," and "Very Low"), and recommendations were formulated (Category A: recommended, Category B: potential with further validation, and Category C: not recommended). RESULTS: Seven studies were included, comprising three specific instruments: the Care Competency Scale for Family Caregivers in Home Palliative Care (CCSHPC) (n = 1), the Caregiver Caregiving Self-Efficacy Scale-Oral Cancer (CSES-OC) (n = 1), and the Caring Ability of Family Caregivers of Patients with Cancer Scale (CAFCPCS) (n = 5). Both the CCSHPC and CAFCPCS received Category B recommendations, demonstrating "adequate" content validity with evidence grades rated "very low" and "low," respectively. The CAFCPCS also shows good structural validity ("moderate") and internal consistency ("low") in some cultural contexts. The CSES-OC is a Category C recommendation, with high-quality evidence indicating "inadequate" criterion validity. CONCLUSION: Few specific instruments exist, and most did not strictly follow COSMIN guidelines. The CAFCPCS is provisionally recommended based on relative evidence superiority rather than complete psychometric validation. Further cross-cultural and localized instrument development is warranted. IMPLICATIONS FOR NURSING PRACTICE: Use well-validated specific instruments to identify strengths and weaknesses in the caregiving competencies of family caregivers of cancer patients, enabling them to deliver high-quality home-based cancer care.

Female

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95 % CI 0.85-0.94; 95 % prediction interval 0.62-0.98), with sensitivity of 0.80 (95 % CI 0.77-0.83) and specificity of 0.87 (95 % CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Determination of 13 per- and polyfluoroalkyl substances in human plasma samples using LC-MS/MS: application to capillary microsamples.

Per- and polyfluoroalkyl substances (PFAS) are chemicals widely applied in industrial processes and highly persistent in the environment, whose extensive use has been linked to adverse health effects. Venous plasma is the conventional matrix for PFAS assessment in blood, and LC-MS/MS is the most used quantification technique. Despite the relevance of this topic, biomonitoring data on human exposure to PFAS in Brazil remain limited. This study validated an LC-MS/MS method for determination of 13 PFAS in human plasma. Blood samples were collected from volunteers by phlebotomy, followed by protein precipitation with acetonitrile containing 1% formic acid (v/v) and solid-phase extraction. Chromatographic separation was achieved on an Acquity UPLC HSS T3 column. The assay was linear over a calibration range of 0.2-20 ng/mL. Intra- and inter-assay precision (CV%) were within the ranges of 2.06-12.0% and 0.25-10.7%, respectively. As for accuracy, results were 89.0-112.9%. Matrix effect ranged from -1.31 to 0.05%. Stability after four freeze/thaw cycles and under autosampler conditions were also confirmed for all analytes. The method was applied to 40 paired venous and capillary plasma samples. Both measures exhibited high correlation (r = 0.926). PFOS was the only compound detected at concentrations ≥0.2 ng/mL (LLOQ) in all samples, with capillary plasma concentrations of 0.85-13.50 ng/mL. In summary, the method showed good validation performance and demonstrated the suitability of capillary plasma samples as an alternative matrix for PFAS quantification.

Humans

Patient-reported outcome measures for depression or anxiety symptoms in patients with cardiovascular disease: A COSMIN systematic review.

BACKGROUND: Depression and anxiety are common in patients with cardiovascular disease (CVD), but the measurement quality of patient-reported outcome measures (PROMs) used in this population remains unclear. This review aimed to evaluate the methodological quality, measurement properties, and certainty of evidence for depression and anxiety PROMs in adults with CVD and to inform instrument selection. METHODS: Following COSMIN and PRISMA guidance, four databases were searched from inception to February 2026. Studies assessing measurement properties of PROMs in adults with CVD were included. Methodological quality was evaluated using the COSMIN Risk of Bias checklist, and certainty of evidence was graded using an adapted GRADE approach. RESULTS: Sixty-six studies assessing 38 PROMs were included, comprising 29 generic and 9 CVD-specific instruments. Six PROMs met COSMIN Category A criteria: Cardiac Depression Scale-Short Form, Patient Health Questionnaire-9, Beck Depression Inventory-II, Hospital Anxiety and Depression Scale, Generalized Anxiety Disorder-7, and Major Depression Inventory. Four instruments were classified as Category C because of insufficient structural validity. Content-validity evidence was largely indeterminate or of limited certainty. Only 24 studies used confirmatory factor analysis or Rasch analysis, and no study assessed measurement error or responsiveness. Cross-cultural validity evidence was scarce. CONCLUSIONS: Six PROMs met Category A criteria, but selection should remain purpose- and context-specific. Particular attention should be given to somatic symptom overlap and intended clinical use. Further validation should prioritize content validity, measurement invariance, responsiveness, measurement error, and clinimetric performance.

Humans

Psychometric Evaluation of the Breast Inflammatory Symptom Severity Index Versions 2 and 3 Among Lactating Women.

OBJECTIVE: To evaluate the psychometric properties of Versions 2 and 3 of the Breast Inflammatory Symptom Severity Index (BISSI). DESIGN: Secondary data analysis of clinical trial data. SETTING: Private physiotherapy practices, a public tertiary hospital, and a community in Melbourne, Australia. PARTICIPANTS: Women more than 7 days after birth with inflammatory conditions of the lactating breast (N = 43). METHODS: We performed confirmatory factor analysis of the BISSI Version 2 to examine item loading, which informed development of the BISSI Version 3 (V3). We assessed convergent validity by comparing total BISSI V3 scores with human milk sodium to potassium ratio (Na+:K+) at Trial Days 1, 3, and 10 using Bland-Altman plots. We compared item-level scores for size of affected area with objective receiver operating characteristic curve analysis to assess discriminant validity for symptom severity and Cronbach's alpha for internal reliability. RESULTS: After confirmatory factor analysis, we removed two items, resulting in a six-item BISSI V3. All retained items demonstrated comparable loading on the overall scale. Limits of agreement for total BISSI V3 scores and item-level scores for size of affected area were acceptable at all time points, with more than 90% of observations falling within 2 standard deviations of the mean difference, supporting convergent validity. Discriminant validity of the BISSI V3 was supported. We found high internal reliability at both time points CONCLUSION: Our findings provide evidence for the validity and reliability of the BISSI V3 and support its continued development and for clinical use of the BISSI V3 and human milk Na+:K+ analysis to enhance management of inflammatory conditions of the lactating breast.

breastfeeding

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

A Dynamic Nomogram to Predict Metabolic Dysfunction-Associated Fatty Liver Disease in Patients with Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) involves multiple metabolic disorders. This study aimed to identify high-risk populations for metabolic dysfunction-associated fatty liver disease (MAFLD) in patients with MetS and to establish a dynamic predictive nomogram. METHODS: A total of 627 patients with MetS from six regions in Zhejiang Province were enrolled and categorized into MAFLD and non-MAFLD groups, then randomly assigned to training and validation sets at a ratio of 7:3. Independent predictors of MAFLD were identified using least absolute shrinkage and selection operator regression and multivariable logistic regression analyses. These predictors were then used to construct a dynamic nomogram. RESULTS: A total of 627 patients with MetS were included in the final analysis, of whom 77.0% (483/627) were diagnosed with MAFLD. Multivariable logistic regression analysis identified body mass index (BMI), waist circumference (WC), total cholesterol (TC), alanine aminotransferase (ALT), MetS-defined dysglycemia, and education level as independent risk factors for MAFLD. MetS-defined dysglycemia showed the highest odds ratio (OR) for MAFLD development [OR = 1.87, 95% confidence interval (CI): 1.07-3.29]. Although the number of MetS components and the metabolic syndrome score were significantly associated with MAFLD in univariate analysis, they were not independently associated with MAFLD in the multivariate model. A dynamic nomogram for predicting MAFLD risk in patients with MetS was developed and internally validated. The area under the receiver operating characteristic curve was 0.834 (95% CI: 0.787-0.880) in the training set and 0.839 (95% CI: 0.771-0.899) in the validation set, indicating strong predictive performance. Bootstrap internal validation demonstrated good agreement between predicted and observed outcomes in calibration curves. Decision curve analysis further indicated favorable clinical applicability of the nomogram. CONCLUSION: BMI, WC, TC, ALT, MetS-defined dysglycemia, and education level are independent risk factors for MAFLD. A dynamic nomogram for predicting MAFLD risk in patients with MetS was successfully developed and validated.

Humans

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Meniscal preservation in the age of biologics: toward a quantitative decision algorithm for personalized repair.

BACKGROUND: Despite advances in arthroscopic repair and biologic augmentation, surgical indication for meniscal tears remains heterogeneous. No standardized framework currently integrates biomechanical, clinical, and biological determinants to guide repair versus resection. PURPOSE: To develop a quantitative decision model-the Meniscal Preservation Score (MPS)-that unifies biomechanical and biological evidence to stratify reparability potential and standardize treatment selection in meniscal surgery. METHODS: A systematic evidence synthesis conducted in accordance with PRISMA 2020 reporting standards of studies published from 2000 to 2025 in PubMed, Embase, and Scopus identified key determinants of meniscal healing. Five consistent predictors-patient age, vascularity, tear morphology, associated pathology, and activity profile-were weighted through a two-round modified Delphi consensus among ten experienced knee surgeons. The resulting 0-9-point MPS was incorporated into a stepwise decision tree linking lesion morphology, biological context, and surgical strategy. Conceptual validation used 50 simulated cases and a retrospective cohort of 45 patients to test agreement between algorithm recommendations and expert surgical decisions. RESULTS: The MPS achieved 86% concordance with expert judgment in simulation and 84% agreement in clinical validation. In this retrospective exploratory cohort, cases in which surgical management was concordant with MPS recommendations demonstrated higher mean IKDC scores at 24 months and lower observed reoperation rates. These findings should be interpreted as associative rather than causal, as treatment allocation was not controlled and discordant cases may have represented inherently more complex pathology. CONCLUSION: The MPS represents an evidence-informed decision-support framework designed to systematize reparability assessment. While exploratory analyses suggest structural coherence with expert reasoning, prospective implementation and external validation are required before clinical adoption as a predictive tool. LEVEL OF EVIDENCE: conceptual model with exploratory validation.

Humans

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n = 38, 74%). Hierarchical clustering (n = 20) and K-means clustering (n = 14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29 709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85) and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

A pragmatic randomized controlled trial of self-directed online writing interventions for posttraumatic stress symptoms in a real-world digital setting.

Background: Public health and other large-scale crises, such as the COVID-19 pandemic, have intensified the global mental health burden, creating unprecedented demand for accessible interventions for posttraumatic stress symptoms (PTSS).Objective: We evaluated the feasibility and effectiveness of two self-directed online writing interventions embedded within China's WeChat ecosystem during the COVID-19 pandemic through a pragmatic randomised controlled trial.Methods: Between December 2021 and August 2022, 1,526 adults were screened for PTSS via a Tencent Medinfo Mini-Program. Eligible participants (n = 211) were randomised to Guided Narrative Technique-Writing (GNT-W, n = 100) or Expressive Writing (EW, n = 111). Both interventions comprised three self-directed daily writing sessions delivered entirely online without human support. Primary outcome was PTSD symptom severity (PTSD Checklist-Short), assessed at baseline, post-intervention, 2-week, and 1-month follow-ups.Results: While initial engagement followed typical digital health patterns (64.5% overall attrition), participants who initiated treatment showed strong adherence (77% completion). Both interventions were associated with significant within-group reductions in PTSS severity (GNT-W: b = -0.43, p = .023, d = -0.43; EW: b = -0.60, p = .001, d = -0.58), with no significant between-group difference (group × time: b = 0.18, p = .48). GNT-W did not confer additional benefit over EW protocol on PTSS severity.Conclusions: Both self-directed writing interventions were associated with within-group reductions in PTSS; without an inactive control condition, however, these changes cannot be firmly attributed to the interventions. GNT-W showed no advantage over the simpler EW protocol. These findings offer preliminary support for embedding scalable, low-barrier writing interventions in widely used digital platforms.Chinese Clinical Trial Registry: ChiCTR2000034836.

Humans

The mighty microproteins: from versatile cellular regulators to precision medicine therapeutics.

Microproteins, are tiny proteins encoded by small open reading frame (sORF), translation of these non-canonical open reading frames (ncORFs) has been implicated in diverse biological processes and diseases. This review summarizes recent developments in the discovery, biogenesis, and functional characterization of microproteins, and their involvement in various disease, with special focus on their roles in cancer, cardiovascular, metabolic, neurodegenerative and immune-related disorders. We emphasize the regulation of key cellular pathways by microproteins, including mitochondrial homeostasis, apoptosis, metabolic reprogramming, and immune signaling, all of which affect disease initiation and progression. Emerging evidence also supports their potential as disease biomarkers and therapeutic candidates for precision medicine. Finally, the review critically discusses the current challenges including discrepancies in microprotein annotation, the limitations of ribosome profiling and proteogenomic approaches, the gap between computationally predicted and experimentally validated microproteins, and the need for rigorous orthogonal validation by means of CRISPR-based genome editing, ribosome release assays, mutational analysis, high-resolution mass spectrometry, and functional studies. Finally, we review recent development of AI-assisted ORF prediction, single-cell translatomics, spatial proteomics, and integrated multi-omics as emerging technologies reshaping. Microprotein discovery and functional annotation. Finally, we discuss the translational potential of microproteins and highlight the remaining challenges to clinical application, including peptide stability, pharmacokinetics, tissue-specific delivery, immunogenicity, and the need for rigorous preclinical and clinical validation. Together, this review provides an updated and critical overview of the rapidly evolving microprotein field and highlights future research priorities for translating these molecules into clinically useful biomarkers and precision therapeutics.

Microproteins

Risk Factors and Predictive Model for Postoperative High Myopia in Children Undergoing Congenital Cataract Surgery With Intraocular Lens Implantation.

PURPOSE: To identify risk factors associated with the development of high myopia following congenital cataract surgery and to establish a robust predictive model. DESIGN: Retrospective clinical cohort study. SUBJECTS: This retrospective study included 106 pediatric patients who underwent congenital cataract surgery with primary IOL implantation (mean follow-up 8.19 years). The model was externally validated in an independent cohort of 72 patients with a mean follow-up of 7.83 years. METHODS: Preoperative and postoperative ocular biometric parameters were collected. Risk factors for postoperative high myopia were analyzed using Cox proportional hazards regression, which served as the basis for model construction. The predictive performance of the model was rigorously evaluated for discrimination and calibration. Discriminative ability was quantified using Harrell's C-index and the area under the receiver operating characteristic curve (AUC). Model calibration was assessed via calibration plots by comparing predicted probabilities with actual observed outcomes. Internal validation was performed using a bootstrapping method (500 iterations) to ensure model stability and adjust for potential overfitting. RESULTS: An initial postoperative refraction of <+0.75D, and a higher IOL Power to Axial length Ratio (IOL/AL ratio) were identified as significant risk factors for the development of postoperative high myopia. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. The predictive model demonstrated robust performance, achieving a C-index of 0.711 (internal validation C-index: 0.713). The area under the receiver operating characteristic curve (AUC) values for predicting high myopia at 5 and 10 years were 0.858 and 0.745, respectively. Furthermore, calibration curves demonstrated excellent agreement between the predicted and observed outcomes throughout the follow-up period. In external validation, the model achieved a C-index of 0.825, 5-year AUC of 0.833, and 10-year AUC of 0.713. CONCLUSIONS: Our analysis established that initial postoperative refraction <+0.75D, and an elevated IOL/AL ratio are key determinants of high myopia risk following surgery. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. This predictive framework provides clinicians with a practical tool to optimize preoperative IOL selection and identify high-risk infants who require vigilant myopia prevention and balanced amblyopia management.

Humans