Search PubMedSearch

SEARCH · Search PubMed

Results for “validation studies”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,498 records · Page 3Linked to original sources

Dynamic evolution of chaperone-mediated autophagy is associated with tumor microenvironment remodeling and prognostic stratification in lung adenocarcinoma: insights from single-cell transcriptomics, ensemble machine learning, and experimental validation.

BACKGROUND: Lung adenocarcinoma (LUAD) shows prognostic heterogeneity, and tumor-node-metastasis (TNM) staging is limited for individualized management. Chaperone-mediated autophagy (CMA) maintains proteostasis, but its role during adenocarcinoma in situ (AIS)-minimally invasive adenocarcinoma (MIA)-invasive adenocarcinoma (IAC) progression remains unclear. METHODS: Single-cell RNA sequencing (scRNA-seq) data from GSE189357 and bulk transcriptomes from The Cancer Genome Atlas (TCGA)-LUAD and Gene Expression Omnibus (GEO) cohorts were integrated. CMA activity, cell-cell communication, weighted gene co-expression network analysis (WGCNA), tumor-normal differential expression, machine-learning survival modeling, tumor microenvironment (TME) features, drug sensitivity, and EPC1 function were analyzed. RESULTS: CMA-high tumor epithelial cells increased from AIS (58.1%) to MIA (65.7%) but declined in IAC (44.4%; p < 0.001). CMA-low cells preferentially received fibroblast-derived extracellular matrix cues. A CMA-negatively correlated module identified 69 core genes. Random survival forest (RSF) performed best among 117 machine-learning combinations (mean concordance index > 0.873). High-risk patients had worse survival across cohorts, and the risk score was independently associated with overall survival (hazard ratio = 16.013, 95% confidence interval: 9.579-26.768, p < 0.001). High-risk tumors showed proliferative activation and M0 macrophage enrichment, whereas low-risk tumors showed stronger immune-related signaling. EPC1 overexpression suppressed malignant phenotypes in A549 cells. CONCLUSION: CMA dynamics are associated with stromal and immune remodeling during LUAD progression. A CMA-based model provides robust prognostic stratification and may offer a basis for future TME-guided studies.

Chaperone-mediated autophagy

Stretched penile length in boys with hypospadias: Population-based analysis using validated nomogram.

BACKGROUND: Hypospadias affects 1 in 200-300 male births. Parents are often concerned about penile adequacy beyond the urethral defect itself, yet few studies have systematically compared stretched penile length (SPL) in hypospadias against population-based reference standards. OBJECTIVE: To evaluate SPL distribution patterns in boys with Types I and II hypospadias and compare them with established normative data. METHODS: The authors studied 876 consecutive boys aged 1-14 years with unoperated Types I (distal) and II (mid-shaft) hypospadias. Two observers independently measured SPL using the validated SPLINT technique. The SPL measurements were compared against age-matched normative data from 1276 Indian children. Exact binomial probability tests were used for percentile distributions, chi-square tests for subtype comparisons and t-tests for mean deviations. RESULTS: The cohort included 479 Type I and 397 Type II cases. SPL distribution showed a marked leftward shift: 71% fell below the 50th percentile (expected 50%, p < 0.001) and 41.5% below the 25th percentile. Lower percentiles were overrepresented, 20.7% were below the 10th percentile and 20.8% in the 10th-25th range. Upper percentiles were depleted: only 7.4% in the 75th-90th range and 1.7% above the 90th percentile (all p < 0.001). Mean SPL was reduced by 6.8% (95% CI: -8.18 to -5.42%) in Type I and 7.5% (95% CI: -9.05 to -5.92%) in Type II. The two subtypes showed no significant distributional difference (&#x3c7;2 = 6.22, p = 0.18), suggesting that meatal position does not predict SPL reduction. CONCLUSIONS: Boys with distal and mid-shaft hypospadias show clinically meaningful SPL reduction that follows a continuous distribution rather than an all-or-none pattern. SPL reduction appears independent of meatal position. These findings support routine SPL assessment using population-specific references and can guide preoperative counselling.

Humans

Validation of the newly introduced Deauville score 5a for patients treated for advanced-stage classic Hodgkin lymphoma.

The Lugano Imaging Committee recently refined the Deauville score (DS), subdividing DS5 into DS5a (>2&#xd7; liver uptake without new lesions) and DS5b (new lesions). We investigated whether this improves prognostic discrimination at interim positron emission tomography (PET) after 2 cycles (PET-2) in patients with advanced-stage classical Hodgkin lymphoma (AS-cHL) treated in recent German Hodgkin Study Group randomized phase 3 trials. The primary analysis cohort was HD18 postamendment standard arms (uniform treatment with 6 cycles of escalated doses of bleomycin, etoposide, doxorubicin, cyclophosphamide, vincristine, procarbazine, and prednisone [eBEACOPP]); sensitivity cohorts were HD18 intention-to-treat and HD21 eBEACOPP and brentuximab vedotin, etoposide, cyclophosphamide, doxorubicin, dacarbazine, and dexamethasone arms. Progression-free survival (PFS) was analyzed by landmark Cox models starting at PET-2. DS5a was infrequent (4%-6% across cohorts; 39/639, 67/1745, 33/568, and 29/560). In the primary cohort, DS5 was associated with inferior PFS vs DS1 to DS3 (hazard ratio [HR], 3.00; 95% confidence interval [CI], 1.25-7.23) and vs DS1 to DS4 (HR, 2.35; 95% CI, 1.01-5.50). Across sensitivity cohorts, DS5a remained adverse compared with DS1 to DS4 (HR range, 2.57-5.47), whereas DS4 according to the new definition did not consistently separate from DS1 to DS3, which is likely a result of PET-adapted treatment. Overall survival trends were concordant, but interpretation is limited by few events. To our knowledge, this is the first prognostic validation of the refined DS in prospectively randomized trial populations. The newly introduced DS5a isolates a small high-risk AS-cHL, which further supports risk assessment and adaptation using quantitative biomarkers from PET. The HD18 and HD21 trials were registered at www.clinicaltrials.gov as NCT00515554 and NCT02661503, respectively.

Humans

Could the preoperative urethral curve be used to predict immediate urinary continence following Retzius-sparing robot-assisted radical prostatectomy? A retrospective multi-center study.

PURPOSE: Immediate urinary continence (UC) recovery following Retzius-sparing robot-assisted radical prostatectomy (RS-RARP) remains highly variable, highlighting the need for reliable preoperative prediction. We aimed to develop and validate models to identify patients likely to achieve immediate UC recovery following RS-RARP. MATERIALS AND METHODS: A total of 580 prostate cancer patients who underwent RS-RARP from four medical centers were assigned to a training set (n=348), an internal validation set (n=103) and an external validation set (n=129). Independent predictors were identified through univariate analysis and LASSO regression. A nomogram was constructed using multivariate logistic regression. Its performance was evaluated with receiver operating characteristic (ROC) curve, calibration curves, and decision curve analysis. RESULTS: Immediate UC recovery was observed in 84.5% (294/348) of patients in the training cohort, 80.6% (83/103) in the internal validation cohort, and 81.4% (105/129) in the external validation cohort, respectively. Multivariate analysis identified membranous urethral length (MUL) (OR=1.23, P=0.029) and urethral curvature (OR=2.84, P<0.001) as independent predictors, while prostate volume (PV) (OR=0.84, P <0.001) as a protective factor. The nomogram integrating MUL, PV, and urethral curvature demonstrated superior predictive accuracy, with an AUC of 0.87 (95% CI, 0.83-0.91) in the training cohort. The bootstrap-corrected calibration slope was 0.96, and the Brier score was 0.08.&#xa0;Calibration curves and decision curve analysis confirmed the predictive accuracy and clinical utility of the nomogram. CONCLUSIONS: Our study introduces a novel quantitative method for assessing urethral curvature. The mpMRI-based model, integrating urethral curvature and prostate spatial configuration, offers enhanced predictive accuracy for postoperative immediate UC recovery.

Humans

Critical insights on the application of the theory of planned behaviour to food handlers' food safety practices.

Foodborne diseases remain a significant public health concern, often linked to unsafe food-handling practices. The Theory of Planned Behaviour (TPB) is widely used to predict and explain food safety behaviours, yet its application in this field has not been systematically and in-depth evaluated. This review evaluated how the TPB has been applied to study food handlers' behaviour, focusing on methodological approaches, use of the TACT (Target, Action, Context, and Time) framework, validity, elicitation studies, and reliability. Seventeen studies were included following a systematic search of four databases (Scopus, Web of Science, Wiley Online Library, and Taylor & Francis Online). Data were extracted on behaviour definition, aim of study, main findings, use of indirect and direct TPB measures, use of elicitation studies, internal consistency, content validation, analytical methods used, and any extensions to the original TPB framework. Key elements related to adherence to core TPB principles and measurement practices were extracted using a Checklist. Most studies used direct measures of TPB constructs, and only a few reported procedures for content validation. Considerable variability was found in the reporting of key measurement and psychometric practices. Five studies fully applied the TACT framework, while nine incorporated additional factors such as knowledge and moral norms. Elicitation studies were conducted in five cases where indirect measures were employed. Analytical approaches were mainly based on multiple linear regression, with limited use of more advanced techniques such as structural equation modeling. Twelve studies reported internal consistency results. Overall, the review highlights opportunities to strengthen methodological practices in future TPB research on food safety. Greater attention to conducting and reporting content validation, full application of the TACT framework, reporting of internal consistency, and consistent inclusion of elicitation studies when using indirect measures may enhance transparency, reinforcing the credibility and trustworthiness of research findings. A major methodological limitation of this review was that screening and data extraction were conducted by a single reviewer and no formal quality or risk-of-bias assessment of the included studies was performed. Despite these limitations, the findings provide practical guidance for the development and validation of TPB-based questionnaires and may support more robust food safety research, interventions, and policy initiatives aimed at improving food handlers' practices.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Vitamin B12 Deficiency in Sickle Cell Disease: Method-Driven Estimates and Systematic Diagnostic Misclassification.

OBJECTIVES: To determine whether the reported 0%-70% prevalence of vitamin B12 deficiency in sickle cell disease (SCD) reflects true population variation or diagnostic misclassification. METHODS: We conducted a PRISMA 2020-compliant systematic review of observational studies (January 1, 2000-May 13, 2026; PROSPERO CRD420251087800) assessing B12 status in SCD. PubMed, AJOL, and Google Scholar were searched with citation tracking and dual screening. Diagnostic validity was assessed across biomarker strategy, analytical platform, thresholds, and confounder control using a proposed context-integrated framework to classify methodological robustness and discordance. RESULTS: Fourteen studies were included (57% high-income; 43% LMIC). The evidence base was dominated by limited diagnostic approaches: 71% used immunoassays, over one-third relied on circulating B12 alone, and functional biomarkers were inconsistently applied without systematic confounder adjustment. Prevalence estimates were strongly influenced by diagnostic methods rather than underlying population biology, ranging from 0% to 70% in single-marker studies (mostly 0%-7.1%, with outliers ~50%-70%) and 6.9%-53% in multi-marker studies. Discordance was substantial and greater in LMIC settings than HIC. CONCLUSION: Current diagnostic approaches in SCD appear method-dependent, generating heterogeneous prevalence estimates with uncertain clinical validity. These findings challenge existing estimates and have implications for clinical practice, research design, and diagnostic equity. TRIAL REGISTRATION: ClinicalTrials.gov identifier: CRD420251087800.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Specific Instruments for Caregiving Competence Among Family Caregivers of Cancer Patients: A COSMIN Systematic Review of Psychometric Properties.

OBJECTIVE: To evaluate and summarize the psychometric properties of specific instruments for caregiving competence among family caregivers of cancer patients. METHODS: Systematically searched eight databases for studies published up to November 2025. The methodological quality and psychometric properties of the instruments were evaluated using COSMIN 2.0. Evidence grades were rated using the modified GRADE system (four grades: "High," "Moderate," "Low," and "Very Low"), and recommendations were formulated (Category A: recommended, Category B: potential with further validation, and Category C: not recommended). RESULTS: Seven studies were included, comprising three specific instruments: the Care Competency Scale for Family Caregivers in Home Palliative Care (CCSHPC) (n = 1), the Caregiver Caregiving Self-Efficacy Scale-Oral Cancer (CSES-OC) (n = 1), and the Caring Ability of Family Caregivers of Patients with Cancer Scale (CAFCPCS) (n = 5). Both the CCSHPC and CAFCPCS received Category B recommendations, demonstrating "adequate" content validity with evidence grades rated "very low" and "low," respectively. The CAFCPCS also shows good structural validity ("moderate") and internal consistency ("low") in some cultural contexts. The CSES-OC is a Category C recommendation, with high-quality evidence indicating "inadequate" criterion validity. CONCLUSION: Few specific instruments exist, and most did not strictly follow COSMIN guidelines. The CAFCPCS is provisionally recommended based on relative evidence superiority rather than complete psychometric validation. Further cross-cultural and localized instrument development is warranted. IMPLICATIONS FOR NURSING PRACTICE: Use well-validated specific instruments to identify strengths and weaknesses in the caregiving competencies of family caregivers of cancer patients, enabling them to deliver high-quality home-based cancer care.

Female

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95&#xa0;% CI 0.85-0.94; 95&#xa0;% prediction interval 0.62-0.98), with sensitivity of 0.80 (95&#xa0;% CI 0.77-0.83) and specificity of 0.87 (95&#xa0;% CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Determination of 13 per- and polyfluoroalkyl substances in human plasma samples using LC-MS/MS: application to capillary microsamples.

Per- and polyfluoroalkyl substances (PFAS) are chemicals widely applied in industrial processes and highly persistent in the environment, whose extensive use has been linked to adverse health effects. Venous plasma is the conventional matrix for PFAS assessment in blood, and LC-MS/MS is the most used quantification technique. Despite the relevance of this topic, biomonitoring data on human exposure to PFAS in Brazil remain limited. This study validated an LC-MS/MS method for determination of 13 PFAS in human plasma. Blood samples were collected from volunteers by phlebotomy, followed by protein precipitation with acetonitrile containing 1% formic acid (v/v) and solid-phase extraction. Chromatographic separation was achieved on an Acquity UPLC HSS T3 column. The assay was linear over a calibration range of 0.2-20&#xa0;ng/mL. Intra- and inter-assay precision (CV%) were within the ranges of 2.06-12.0% and 0.25-10.7%, respectively. As for accuracy, results were 89.0-112.9%. Matrix effect ranged from -1.31 to 0.05%. Stability after four freeze/thaw cycles and under autosampler conditions were also confirmed for all analytes. The method was applied to 40 paired venous and capillary plasma samples. Both measures exhibited high correlation (r&#xa0;=&#xa0;0.926). PFOS was the only compound detected at concentrations &#x2265;0.2&#xa0;ng/mL (LLOQ) in all samples, with capillary plasma concentrations of 0.85-13.50&#xa0;ng/mL. In summary, the method showed good validation performance and demonstrated the suitability of capillary plasma samples as an alternative matrix for PFAS quantification.

Humans

Patient-reported outcome measures for depression or anxiety symptoms in patients with cardiovascular disease: A COSMIN systematic review.

BACKGROUND: Depression and anxiety are common in patients with cardiovascular disease (CVD), but the measurement quality of patient-reported outcome measures (PROMs) used in this population remains unclear. This review aimed to evaluate the methodological quality, measurement properties, and certainty of evidence for depression and anxiety PROMs in adults with CVD and to inform instrument selection. METHODS: Following COSMIN and PRISMA guidance, four databases were searched from inception to February 2026. Studies assessing measurement properties of PROMs in adults with CVD were included. Methodological quality was evaluated using the COSMIN Risk of Bias checklist, and certainty of evidence was graded using an adapted GRADE approach. RESULTS: Sixty-six studies assessing 38 PROMs were included, comprising 29 generic and 9 CVD-specific instruments. Six PROMs met COSMIN Category A criteria: Cardiac Depression Scale-Short Form, Patient Health Questionnaire-9, Beck Depression Inventory-II, Hospital Anxiety and Depression Scale, Generalized Anxiety Disorder-7, and Major Depression Inventory. Four instruments were classified as Category C because of insufficient structural validity. Content-validity evidence was largely indeterminate or of limited certainty. Only 24 studies used confirmatory factor analysis or Rasch analysis, and no study assessed measurement error or responsiveness. Cross-cultural validity evidence was scarce. CONCLUSIONS: Six PROMs met Category A criteria, but selection should remain purpose- and context-specific. Particular attention should be given to somatic symptom overlap and intended clinical use. Further validation should prioritize content validity, measurement invariance, responsiveness, measurement error, and clinimetric performance.

Humans

Psychometric Evaluation of the Breast Inflammatory Symptom Severity Index Versions 2 and 3 Among Lactating Women.

OBJECTIVE: To evaluate the psychometric properties of Versions 2 and 3 of the Breast Inflammatory Symptom Severity Index (BISSI). DESIGN: Secondary data analysis of clinical trial data. SETTING: Private physiotherapy practices, a public tertiary hospital, and a community in Melbourne, Australia. PARTICIPANTS: Women more than 7 days after birth with inflammatory conditions of the lactating breast (N = 43). METHODS: We performed confirmatory factor analysis of the BISSI Version 2 to examine item loading, which informed development of the BISSI Version 3 (V3). We assessed convergent validity by comparing total BISSI V3 scores with human milk sodium to potassium ratio (Na+:K+) at Trial Days 1, 3, and 10 using Bland-Altman plots. We compared item-level scores for size of affected area with objective receiver operating characteristic curve analysis to assess discriminant validity for symptom severity and Cronbach's alpha for internal reliability. RESULTS: After confirmatory factor analysis, we removed two items, resulting in a six-item BISSI V3. All retained items demonstrated comparable loading on the overall scale. Limits of agreement for total BISSI V3 scores and item-level scores for size of affected area were acceptable at all time points, with more than 90% of observations falling within 2 standard deviations of the mean difference, supporting convergent validity. Discriminant validity of the BISSI V3 was supported. We found high internal reliability at both time points CONCLUSION: Our findings provide evidence for the validity and reliability of the BISSI V3 and support its continued development and for clinical use of the BISSI V3 and human milk Na+:K+ analysis to enhance management of inflammatory conditions of the lactating breast.

breastfeeding

The Soundtrack of Everyday Life: Real-world Music Listening Habits of Adult Cochlear Implant Users.

OBJECTIVE: Characterize real-world patterns of music listening and reward sensitivity among adult cochlear implant (CI) users compared with normal-hearing (NH) listeners. STUDY DESIGN: Cross-sectional observational study. SETTING: Online. PATIENTS: Adults (&#x2265;18&#xa0;y) with a CI or NH who used a music-streaming platform as their primary listening method. INTERVENTIONS: None. MAIN OUTCOME MEASURES: Objective measures included platform-derived audio features (acousticness, danceability, energy, tempo, and valence), listening volume, unique-song ratio, and decade preferences. Self-reported measures included listening habits and the Barcelona Music Reward Questionnaire (BMRQ). Group comparisons used ANCOVAs and mixed-effect models adjusting for age and gender; within-CI analyses compared prelingual versus postlingual deafness. RESULTS: Among 16 CI users (69% male, 39.0&#xb1;15.6&#xa0;y) and 29 NH listeners (38% male, 32.2&#xb1;11.3&#xa0;y), CI users demonstrated a higher unique-song ratio (&#x3b2;=-0.157, CI as reference; 95%CI [-0.281, -0.034]; P =0.014) and stronger preference for older music (Pillai trace=0.537; F8,35 =5.07; P <0.001), adjusting for age. No significant group differences were observed in weekly listening time, listening volume, audio features, or BMRQ scores (total and subscores). Equivalence was confirmed for Emotion Evocation and Sensory-Motor subscales. There were no statistically significant differences in listening environments after Holm correction. No statistically detectable differences were observed between pre- and postlingually deafened CI users. CONCLUSIONS: Musically active CI users showed no significant differences in listening volume or overall music-reward sensitivity compared with NH peers, but demonstrated higher unique-song ratios and a bias toward older music. Findings highlight the value of ecologically valid data in understanding real-world music experiences among CI users.

Adult

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

A Dynamic Nomogram to Predict Metabolic Dysfunction-Associated Fatty Liver Disease in Patients with Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) involves multiple metabolic disorders. This study aimed to identify high-risk populations for metabolic dysfunction-associated fatty liver disease (MAFLD) in patients with MetS and to establish a dynamic predictive nomogram. METHODS: A total of 627 patients with MetS from six regions in Zhejiang Province were enrolled and categorized into MAFLD and non-MAFLD groups, then randomly assigned to training and validation sets at a ratio of 7:3. Independent predictors of MAFLD were identified using least absolute shrinkage and selection operator regression and multivariable logistic regression analyses. These predictors were then used to construct a dynamic nomogram. RESULTS: A total of 627 patients with MetS were included in the final analysis, of whom 77.0% (483/627) were diagnosed with MAFLD. Multivariable logistic regression analysis identified body mass index (BMI), waist circumference (WC), total cholesterol (TC), alanine aminotransferase (ALT), MetS-defined dysglycemia, and education level as independent risk factors for MAFLD. MetS-defined dysglycemia showed the highest odds ratio (OR) for MAFLD development [OR = 1.87, 95% confidence interval (CI): 1.07-3.29]. Although the number of MetS components and the metabolic syndrome score were significantly associated with MAFLD in univariate analysis, they were not independently associated with MAFLD in the multivariate model. A dynamic nomogram for predicting MAFLD risk in patients with MetS was developed and internally validated. The area under the receiver operating characteristic curve was 0.834 (95% CI: 0.787-0.880) in the training set and 0.839 (95% CI: 0.771-0.899) in the validation set, indicating strong predictive performance. Bootstrap internal validation demonstrated good agreement between predicted and observed outcomes in calibration curves. Decision curve analysis further indicated favorable clinical applicability of the nomogram. CONCLUSION: BMI, WC, TC, ALT, MetS-defined dysglycemia, and education level are independent risk factors for MAFLD. A dynamic nomogram for predicting MAFLD risk in patients with MetS was successfully developed and validated.

Humans

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans