Search PubMedSearch

SEARCH · Search PubMed

Results for “ecological validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

447 records · Page 7Linked to original sources

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Effects of a situational interest-based rotational fitness training program on positive affect in university students: A randomized controlled trial.

BACKGROUND: To evaluate whether a 12-week situational interest-based rotational fitness program improved positive emotional states in university students, to compare module-specific effects, and to examine whether the intervention moderated the heart rate-emotion association. METHODS: In this 2-arm randomized controlled trial, 560 freshmen from 10 Chinese universities were assigned to an experimental (n&#x2005;=&#x2005;280) or control group (n&#x2005;=&#x2005;280). The experimental group completed weekly 90-minute sessions covering strength, endurance, speed, and agility modules; controls received regular physical education. Positive emotional states were assessed using the positive well-being subscale of the Subjective Exercise Experiences Scale through ecological momentary assessment at 3 points per session, yielding over 14,000 observations. Linear mixed-effects models tested group, time, group&#x2005;&#xd7;&#x2005;time, heart rate, and heart rate&#x2005;&#xd7;&#x2005;group effects. Assessors and analysts were blinded. Ethical approval for the trial was granted by the Institutional Review Board of Capital University of Physical Education and Sports (2025A055). RESULTS: Improvement was greater in the experimental group than in controls (EMM change: 1.843 vs 0.505; difference&#x2005;=&#x2005;1.338, 95% confidence interval [CI]: 1.275-1.400, P&#x2005;<&#x2005;.001). At week 12, the raw between-group difference was 1.85 (d&#x2005;=&#x2005;2.23), and all group&#x2005;&#xd7;&#x2005;week interactions from weeks 2 to 12 were significant. Heart rate positively predicted emotional states in controls (B&#x2005;=&#x2005;0.005, 95% CI: 0.003-0.007), but this association was attenuated in the experimental group (interaction B&#x2005;=&#x2005;-0.008, 95% CI: -0.011 to -0.005). Strength and agility showed significant improvements, while endurance produced consistently higher emotional states across assessment points. Heart rate moderation was significant for strength and endurance, but not speed or agility. CONCLUSION: The rotational program improved positive emotional states more than regular physical education, although effects varied by module. Strength showed the most robust benefit. The program also weakened, but did not eliminate, the heart rate-emotion association during more demanding activities.

Humans

Specific Instruments for Caregiving Competence Among Family Caregivers of Cancer Patients: A COSMIN Systematic Review of Psychometric Properties.

OBJECTIVE: To evaluate and summarize the psychometric properties of specific instruments for caregiving competence among family caregivers of cancer patients. METHODS: Systematically searched eight databases for studies published up to November 2025. The methodological quality and psychometric properties of the instruments were evaluated using COSMIN 2.0. Evidence grades were rated using the modified GRADE system (four grades: "High," "Moderate," "Low," and "Very Low"), and recommendations were formulated (Category A: recommended, Category B: potential with further validation, and Category C: not recommended). RESULTS: Seven studies were included, comprising three specific instruments: the Care Competency Scale for Family Caregivers in Home Palliative Care (CCSHPC) (n = 1), the Caregiver Caregiving Self-Efficacy Scale-Oral Cancer (CSES-OC) (n = 1), and the Caring Ability of Family Caregivers of Patients with Cancer Scale (CAFCPCS) (n = 5). Both the CCSHPC and CAFCPCS received Category B recommendations, demonstrating "adequate" content validity with evidence grades rated "very low" and "low," respectively. The CAFCPCS also shows good structural validity ("moderate") and internal consistency ("low") in some cultural contexts. The CSES-OC is a Category C recommendation, with high-quality evidence indicating "inadequate" criterion validity. CONCLUSION: Few specific instruments exist, and most did not strictly follow COSMIN guidelines. The CAFCPCS is provisionally recommended based on relative evidence superiority rather than complete psychometric validation. Further cross-cultural and localized instrument development is warranted. IMPLICATIONS FOR NURSING PRACTICE: Use well-validated specific instruments to identify strengths and weaknesses in the caregiving competencies of family caregivers of cancer patients, enabling them to deliver high-quality home-based cancer care.

Female

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95&#xa0;% CI 0.85-0.94; 95&#xa0;% prediction interval 0.62-0.98), with sensitivity of 0.80 (95&#xa0;% CI 0.77-0.83) and specificity of 0.87 (95&#xa0;% CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Determination of 13 per- and polyfluoroalkyl substances in human plasma samples using LC-MS/MS: application to capillary microsamples.

Per- and polyfluoroalkyl substances (PFAS) are chemicals widely applied in industrial processes and highly persistent in the environment, whose extensive use has been linked to adverse health effects. Venous plasma is the conventional matrix for PFAS assessment in blood, and LC-MS/MS is the most used quantification technique. Despite the relevance of this topic, biomonitoring data on human exposure to PFAS in Brazil remain limited. This study validated an LC-MS/MS method for determination of 13 PFAS in human plasma. Blood samples were collected from volunteers by phlebotomy, followed by protein precipitation with acetonitrile containing 1% formic acid (v/v) and solid-phase extraction. Chromatographic separation was achieved on an Acquity UPLC HSS T3 column. The assay was linear over a calibration range of 0.2-20&#xa0;ng/mL. Intra- and inter-assay precision (CV%) were within the ranges of 2.06-12.0% and 0.25-10.7%, respectively. As for accuracy, results were 89.0-112.9%. Matrix effect ranged from -1.31 to 0.05%. Stability after four freeze/thaw cycles and under autosampler conditions were also confirmed for all analytes. The method was applied to 40 paired venous and capillary plasma samples. Both measures exhibited high correlation (r&#xa0;=&#xa0;0.926). PFOS was the only compound detected at concentrations &#x2265;0.2&#xa0;ng/mL (LLOQ) in all samples, with capillary plasma concentrations of 0.85-13.50&#xa0;ng/mL. In summary, the method showed good validation performance and demonstrated the suitability of capillary plasma samples as an alternative matrix for PFAS quantification.

Humans

Patient-reported outcome measures for depression or anxiety symptoms in patients with cardiovascular disease: A COSMIN systematic review.

BACKGROUND: Depression and anxiety are common in patients with cardiovascular disease (CVD), but the measurement quality of patient-reported outcome measures (PROMs) used in this population remains unclear. This review aimed to evaluate the methodological quality, measurement properties, and certainty of evidence for depression and anxiety PROMs in adults with CVD and to inform instrument selection. METHODS: Following COSMIN and PRISMA guidance, four databases were searched from inception to February 2026. Studies assessing measurement properties of PROMs in adults with CVD were included. Methodological quality was evaluated using the COSMIN Risk of Bias checklist, and certainty of evidence was graded using an adapted GRADE approach. RESULTS: Sixty-six studies assessing 38 PROMs were included, comprising 29 generic and 9 CVD-specific instruments. Six PROMs met COSMIN Category A criteria: Cardiac Depression Scale-Short Form, Patient Health Questionnaire-9, Beck Depression Inventory-II, Hospital Anxiety and Depression Scale, Generalized Anxiety Disorder-7, and Major Depression Inventory. Four instruments were classified as Category C because of insufficient structural validity. Content-validity evidence was largely indeterminate or of limited certainty. Only 24 studies used confirmatory factor analysis or Rasch analysis, and no study assessed measurement error or responsiveness. Cross-cultural validity evidence was scarce. CONCLUSIONS: Six PROMs met Category A criteria, but selection should remain purpose- and context-specific. Particular attention should be given to somatic symptom overlap and intended clinical use. Further validation should prioritize content validity, measurement invariance, responsiveness, measurement error, and clinimetric performance.

Humans

Psychometric Evaluation of the Breast Inflammatory Symptom Severity Index Versions 2 and 3 Among Lactating Women.

OBJECTIVE: To evaluate the psychometric properties of Versions 2 and 3 of the Breast Inflammatory Symptom Severity Index (BISSI). DESIGN: Secondary data analysis of clinical trial data. SETTING: Private physiotherapy practices, a public tertiary hospital, and a community in Melbourne, Australia. PARTICIPANTS: Women more than 7 days after birth with inflammatory conditions of the lactating breast (N = 43). METHODS: We performed confirmatory factor analysis of the BISSI Version 2 to examine item loading, which informed development of the BISSI Version 3 (V3). We assessed convergent validity by comparing total BISSI V3 scores with human milk sodium to potassium ratio (Na+:K+) at Trial Days 1, 3, and 10 using Bland-Altman plots. We compared item-level scores for size of affected area with objective receiver operating characteristic curve analysis to assess discriminant validity for symptom severity and Cronbach's alpha for internal reliability. RESULTS: After confirmatory factor analysis, we removed two items, resulting in a six-item BISSI V3. All retained items demonstrated comparable loading on the overall scale. Limits of agreement for total BISSI V3 scores and item-level scores for size of affected area were acceptable at all time points, with more than 90% of observations falling within 2 standard deviations of the mean difference, supporting convergent validity. Discriminant validity of the BISSI V3 was supported. We found high internal reliability at both time points CONCLUSION: Our findings provide evidence for the validity and reliability of the BISSI V3 and support its continued development and for clinical use of the BISSI V3 and human milk Na+:K+ analysis to enhance management of inflammatory conditions of the lactating breast.

breastfeeding

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

A Dynamic Nomogram to Predict Metabolic Dysfunction-Associated Fatty Liver Disease in Patients with Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) involves multiple metabolic disorders. This study aimed to identify high-risk populations for metabolic dysfunction-associated fatty liver disease (MAFLD) in patients with MetS and to establish a dynamic predictive nomogram. METHODS: A total of 627 patients with MetS from six regions in Zhejiang Province were enrolled and categorized into MAFLD and non-MAFLD groups, then randomly assigned to training and validation sets at a ratio of 7:3. Independent predictors of MAFLD were identified using least absolute shrinkage and selection operator regression and multivariable logistic regression analyses. These predictors were then used to construct a dynamic nomogram. RESULTS: A total of 627 patients with MetS were included in the final analysis, of whom 77.0% (483/627) were diagnosed with MAFLD. Multivariable logistic regression analysis identified body mass index (BMI), waist circumference (WC), total cholesterol (TC), alanine aminotransferase (ALT), MetS-defined dysglycemia, and education level as independent risk factors for MAFLD. MetS-defined dysglycemia showed the highest odds ratio (OR) for MAFLD development [OR = 1.87, 95% confidence interval (CI): 1.07-3.29]. Although the number of MetS components and the metabolic syndrome score were significantly associated with MAFLD in univariate analysis, they were not independently associated with MAFLD in the multivariate model. A dynamic nomogram for predicting MAFLD risk in patients with MetS was developed and internally validated. The area under the receiver operating characteristic curve was 0.834 (95% CI: 0.787-0.880) in the training set and 0.839 (95% CI: 0.771-0.899) in the validation set, indicating strong predictive performance. Bootstrap internal validation demonstrated good agreement between predicted and observed outcomes in calibration curves. Decision curve analysis further indicated favorable clinical applicability of the nomogram. CONCLUSION: BMI, WC, TC, ALT, MetS-defined dysglycemia, and education level are independent risk factors for MAFLD. A dynamic nomogram for predicting MAFLD risk in patients with MetS was successfully developed and validated.

Humans

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Temporal proteomic analysis reveals a three-phase adaptation strategy in Phytophthora cinnamomi during salinity stress.

Phytophthora cinnamomi, a highly invasive hemibiotrophic oomycete, threatens global agriculture, forestry, and native ecosystems. Although drought and temperature effects on P. cinnamomi-host interactions are well studied, current knowledge of abiotic stress responses in P. cinnamomi remains largely centered on infection and phytopathology, with limited molecular insight into the pathogen's direct response to salinity independent of its host. To address this gap, we combined growth assays, time-resolved proteomics, and network analysis to define how P. cinnamomi responds and adapts to salinity exposure. Growth assays showed that NaCl-modified agar enhanced mycelial expansion in a concentration-dependent manner, with 100&#xa0;mM NaCl significantly increasing growth at 48, 72, and 96&#xa0;h compared with controls, while 50&#xa0;mM NaCl remained comparable to control conditions. Temporal proteomic analysis of 100&#xa0;mM NaCl treatment at 0, 1, 6, 12, and 24&#xa0;h post treatment revealed dynamic shifts in protein abundance. Early induction of ROS (Reactive Oxygen Species)-detoxifying enzymes, including glutathione S-transferases and peroxidases, was consistent with ROS-specific staining assays. Network analysis identified modules enriched for redox regulation, ATP generation, ion transport, and translational control, highlighting multi-layered adaptation to elevated NaCl levels. Notably, clusters of conserved hypothetical proteins were strongly upregulated, indicating unexplored stress tolerance components in Phytophthora species. Here, we propose that P. cinnamomi rapidly activates a three-phase strategy involving metabolism readjustments, redox defenses, and cellular structure alterations under salinity conditions. With increasing soil salinization due to climate change, our study provides first mechanistic insights into P. cinnamomi's adaptive plasticity and ecological resilience to abiotic stress. SIGNIFICANCE: This study represents the first temporal proteomic analysis of salinity stress adaptation in Phytophthora cinnamomi, revealing a sophisticated three-phase adaptation strategy. This research fundamentally advances our understanding of how this globally destructive plant pathogen, P. cinnamomi, maintains environmental resilience. Our findings reveal proteome remodelling as a mechanistic framework for understanding stress tolerance in oomycetes, a group of microorganisms responsible for some of the world's most destructive agricultural and forest diseases. Our results show proteins involved in emergency damage control through metabolic recalibration to sustained adaptation. These findings have relevance for predicting pathogen behavior under climate change scenarios, where increasing soil salinity threatens agricultural productivity while simultaneously enhancing pathogen survival and virulence. Understanding how P. cinnamomi responds to prolonged salinity exposure may inform targeted biocontrol strategies and improve predictive models of disease pressure in salt-affected agricultural regions. The temporal analysis framework we present offers a broadly applicable approach for understanding microbial stress adaptation, with implications extending beyond plant pathology to environmental microbiology and biotechnology applications where stress tolerance is paramount.

Phytophthora

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Meniscal preservation in the age of biologics: toward a quantitative decision algorithm for personalized repair.

BACKGROUND: Despite advances in arthroscopic repair and biologic augmentation, surgical indication for meniscal tears remains heterogeneous. No standardized framework currently integrates biomechanical, clinical, and biological determinants to guide repair versus resection. PURPOSE: To develop a quantitative decision model-the Meniscal Preservation Score (MPS)-that unifies biomechanical and biological evidence to stratify reparability potential and standardize treatment selection in meniscal surgery. METHODS: A systematic evidence synthesis conducted in accordance with PRISMA 2020 reporting standards of studies published from 2000 to 2025 in PubMed, Embase, and Scopus identified key determinants of meniscal healing. Five consistent predictors-patient age, vascularity, tear morphology, associated pathology, and activity profile-were weighted through a two-round modified Delphi consensus among ten experienced knee surgeons. The resulting 0-9-point MPS was incorporated into a stepwise decision tree linking lesion morphology, biological context, and surgical strategy. Conceptual validation used 50 simulated cases and a retrospective cohort of 45 patients to test agreement between algorithm recommendations and expert surgical decisions. RESULTS: The MPS achieved 86% concordance with expert judgment in simulation and 84% agreement in clinical validation. In this retrospective exploratory cohort, cases in which surgical management was concordant with MPS recommendations demonstrated higher mean IKDC scores at 24&#xa0;months and lower observed reoperation rates. These findings should be interpreted as associative rather than causal, as treatment allocation was not controlled and discordant cases may have represented inherently more complex pathology. CONCLUSION: The MPS represents an evidence-informed decision-support framework designed to systematize reparability assessment. While exploratory analyses suggest structural coherence with expert reasoning, prospective implementation and external validation are required before clinical adoption as a predictive tool. LEVEL OF EVIDENCE: conceptual model with exploratory validation.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Urinary Small Extracellular Vesicle DNA as a Biomarker for the Non-Invasive Diagnosis of Bladder Cancer.

Existing diagnostic technologies for bladder cancer (BC) suffer from low sensitivity, low specificity, or a lack of validation. Therefore, validated, non-invasive diagnostic biomarkers with high sensitivity and specificity for early detection of BC are needed to complement and improve upon the limitations of existing diagnostic methods. We used low-pass whole genome sequencing (LP-WGS) technology to detect copy number variations (CNVs) in small extracellular vesicle (sEV) DNA isolated from urine samples of patients. Based on these results, we constructed and validated a diagnostic model to differentiate between benign and malignant bladder lesions. We conducted a receiver operating characteristic analysis and calculated the area under the curve (AUC) to evaluate the performance of the diagnostic model. The urine sEV-DNA LP-WGS data revealed CNV differences between benign and malignant samples. The diagnostic model achieved an AUC of 0.953, a sensitivity of 86.7%, and a specificity of 100% in the training cohort and an AUC of 0.985, a sensitivity of 90%, and a specificity of 100% in the validation cohort. Even at the lowest coverage depth of 0.01X, the performance of the diagnostic model remained relatively robust. Notably, the performance of this diagnostic model surpassed that of the biomarker neuron-specific enolase (sensitivity: 85.7% vs. 64.3%; specificity: 100% vs. 87.5%) and urinary cytology (sensitivity: 100% vs. 66.7%; specificity: 100% vs. 94.1%). Our study demonstrates that urine sEV-DNA exhibits high discriminatory power in distinguishing between benign and malignant bladder lesions, making it a promising tool for auxiliary diagnosis of BC.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

The mighty microproteins: from versatile cellular regulators to precision medicine therapeutics.

Microproteins, are tiny proteins encoded by small open reading frame (sORF), translation of these non-canonical open reading frames (ncORFs) has been implicated in diverse biological processes and diseases. This review summarizes recent developments in the discovery, biogenesis, and functional characterization of microproteins, and their involvement in various disease, with special focus on their roles in cancer, cardiovascular, metabolic, neurodegenerative and immune-related disorders. We emphasize the regulation of key cellular pathways by microproteins, including mitochondrial homeostasis, apoptosis, metabolic reprogramming, and immune signaling, all of which affect disease initiation and progression. Emerging evidence also supports their potential as disease biomarkers and therapeutic candidates for precision medicine. Finally, the review critically discusses the current challenges including discrepancies in microprotein annotation, the limitations of ribosome profiling and proteogenomic approaches, the gap between computationally predicted and experimentally validated microproteins, and the need for rigorous orthogonal validation by means of CRISPR-based genome editing, ribosome release assays, mutational analysis, high-resolution mass spectrometry, and functional studies. Finally, we review recent development of AI-assisted ORF prediction, single-cell translatomics, spatial proteomics, and integrated multi-omics as emerging technologies reshaping. Microprotein discovery and functional annotation. Finally, we discuss the translational potential of microproteins and highlight the remaining challenges to clinical application, including peptide stability, pharmacokinetics, tissue-specific delivery, immunogenicity, and the need for rigorous preclinical and clinical validation. Together, this review provides an updated and critical overview of the rapidly evolving microprotein field and highlights future research priorities for translating these molecules into clinically useful biomarkers and precision therapeutics.

Microproteins