Search PubMedSearch

SEARCH · Search PubMed

Results for “performance measurement”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,261 recordsLinked to original sources

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results

Clinical performance of monolithic and veneered zirconia three-unit posterior FDPs: A five-year multicenter randomized controlled trial.

AIM: This randomized controlled clinical study compared monolithic, partially veneered, and fully veneered zirconia FDPs over a 5-year period with respect to survival, technical and biological complications, and patient-reported outcome measures (PROMs). MATERIALS AND METHODS: Sixty-four patients requiring three-unit posterior FDPs were randomly allocated to monolithic (MONO-FDP), partially veneered (PV-FDP), or fully veneered (FV-FDP) groups. All FDPs were fabricated from 4 mol% Y₂O₃ partially stabilized zirconia (4Y-TZP). Follow-up examinations were conducted at baseline, 1, 3, and 5 years. Technical parameters were evaluated using modified USPHS criteria. Periodontal measurements (PPD, BOP, PI) and patient satisfaction were assessed at all time points. RESULTS: A total of 63 FDPs were evaluated at baseline, 57 at 3 years, and 54 at 5 years. Survival at 5 years was 94.7% for MONO-FDPs, 100% for FV-FDPs and 100% for PV-FDPs. Technical complications occurred exclusively in PV-FDP (33%) and FV-FDP (38%) groups and consisted of minor, polishable chipping; no chipping or fractures were recorded in MONO-FDPs (0%) with statistically significant difference between MONO-FDP and the other two groups (p ≤ 0.014). Biological parameters remained stable across all groups, with no significant differences in PPD, BOP, or PI. PV-FDPs and FV-FDPs tended to receive more favorable professional color ratings, whereas PROMs were similar among the groups. CONCLUSIONS: All three FDP designs-monolithic, partially veneered, and fully veneered-fabricated from 4Y-TZP zirconia demonstrated excellent 5-year clinical performance. Technical complications were limited to PV- and FV-FDPs and consisted of minor chipping. Biological outcomes and patient satisfaction were similar across groups. CLINICAL SIGNIFICANCE: The five-year outcomes suggest that veneered, partially veneered, and monolithic FDPs can be used with high clinical reliability; however, restorations incorporating veneering ceramic may present an elevated risk of ceramic chipping.

Humans

Temporal redistribution of control reveals age-related differences in task switching at the level of preparation.

Task-switching studies often report minimal age-related differences in switch costs, leading to the conclusion that switching-related control processes are relatively preserved in aging. However, this conclusion is based on paradigms that confound preparatory and execution processes. This study examined whether age-related differences in semantic task-set reconfiguration may be underestimated due to this confound. In Experiment 1 (36 young and 30 older adults), participants performed an externally paced task-switching paradigm without control over preparation. In Experiment 2 (28 young and 28 older adults), a self-paced paradigm allowed participants to initiate stimulus onset, enabling measurement of preparation time. Across both experiments, reaction time (RT) and error rate (ER) showed reliable age effects but no interactions between age and condition, whereas switching-related condition effects varied across measures and experiments. The expression of switching-related costs differed across measures and task structures. Local switch costs were expressed in ER in Experiment 1 but in RT in Experiment 2. Global switch costs (all-switch vs. all-repeat) were observed in execution measures only in Experiment 1. In Experiment 2, preparation time showed reliable mixing, local, and global switching effects, with age-related amplification emerging specifically for global switching. These findings indicate that switching-related costs are redistributed across processing stages and behavioral measures. The results suggest that age-related modulation of semantic task-set reconfiguration may emerge more clearly during preparation than task execution, particularly under continuous switching demands. Preparation time is interpreted cautiously as reflecting participant-regulated preparatory processes rather than a pure measure of preparation efficiency.

Humans

Impact of neoprene wetsuits on lung volumes and work of breathing: implications for military diver safety and performance.

INTRODUCTION: Neoprene wetsuits may impose mechanical constraints on the chest wall, potentially altering respiratory function. This study investigated the impact of neoprene wetsuits on lung volumes, airway mechanics, and work of breathing (WOB) in healthy male divers. METHODS: A randomised crossover trial was conducted with 31 male divers at the Royal Netherlands Navy Diving Medical Centre. Participants underwent pulmonary function testing, including spirometry, body plethysmography, the forced oscillation technique (FOT), and diffusion capacity measurements, both with and without a hoodless standardised 5 mm neoprene full body wetsuit with a neoprene neck seal. Primary outcomes included changes in forced vital capacity (FVC), functional residual capacity (FRC), airway resistance (Raw), reactance (Xrs), and WOB. RESULTS: Wearing a neoprene wetsuit led to statistically significant reductions in FVC (2.8%, P < 0.05), forced expiration in one second (2.9%, P < 0.05), FRC (4.0%, P < 0.05), and expiratory reserve volume (10.9%, P < 0.05), alongside increases in inspiratory capacity and tidal volume. Raw increased significantly (P < 0.05), while the FOT revealed altered airway mechanics, evidenced by increased Xrs at multiple frequencies (P < 0.05). Diffusion capacity remained unchanged, suggesting preserved alveolar-capillary function. CONCLUSIONS: Neoprene wetsuits induce mechanically restrictive effects on the chest wall, reducing static and dynamic lung volumes and increasing WOB. While these changes may not be clinically relevant at rest, their impact needs to be determined during strenuous or prolonged dives, particularly when combined with other equipment that limits thorax excursions. Future research should explore the effects of the military 5 mm wetsuit under immersed conditions to better understand their operational impact on diver performance and safety.

Male

Immersive virtual reality-assisted anatomy training improves endotracheal intubation performance in simulation: a randomized controlled trial among Chinese non-anesthesiology residents.

INTRODUCTION: This study aimed to compare immersive virtual reality (IVR)-assisted versus conventional anatomy training for teaching endotracheal intubation (ETI) to novice non-anesthesiology residents enrolled in China's Standardized Residency Training program. METHODS: A total of 90 non-anesthesiology residents without prior ETI experience were randomly assigned to either an IVR group receiving IVR-assisted anatomy training (n&#x2009;=&#x2009;45) or a control group receiving conventional anatomy training (n&#x2009;=&#x2009;45). All participants underwent a standardized teaching protocol. The primary endpoint was residents' ETI performance on a simulator, assessed using both the Global Rating Scale (GRS) and a task-specific checklist. The secondary endpoints included changes in written multiple-choice question (MCQ) scores and residents' evaluations of the course. RESULTS: In practical ETI assessments on a manikin, the IVR group achieved significantly higher scores on the task-specific checklist than the control group (90.34&#x2009;&#xb1;&#x2009;2.89 vs. 87.20&#x2009;&#xb1;&#x2009;3.29; p&#x2009;<&#x2009;0.001), whereas GRS scores were comparable between groups. Both groups showed significant post-training improvement in knowledge scores (p&#x2009;<&#x2009;0.001), with the IVR group showing a greater gain in theoretical knowledge (54.0% vs. 36.3%; p&#x2009;<&#x2009;0.001). Participants in the IVR group also expressed a stronger preference for their training method (80.8%) and reported higher levels of motivation, confidence, and enjoyment (all p&#x2009;<&#x2009;0.05). CONCLUSION: IVR-assisted anatomy training enhances the effectiveness of ETI training for novice non-anesthesiology residents, offering an interactive, engaging, and reproducible approach within China's Standardized Residency Training framework.

Humans

Proprioception Training and Surrogate Outcomes: A Systematic Review of Definitions, Measures, and Effectiveness Claims.

BACKGROUND: "Proprioception training" is widely advocated in rehabilitation and sports practice, yet the term encompasses heterogeneous constructs, interventions, and outcomes. Many trials infer proprioceptive benefits from surrogate outcomes (balance, strength, or pain) rather than direct psychophysical indices. OBJECTIVE: We aimed to examine how proprioception is defined and measured, and how improvement is claimed, in randomized controlled trials. METHODS: Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020, PubMed, Scopus, and Web of Science were searched to October 2025. Eligible randomized controlled trials explicitly described interventions as "proprioceptive" or "sensorimotor training" and reported at least one proprioceptive outcome, either direct (e.g., joint position reproduction, threshold to detection of passive motion, active movement extent discrimination) or indirect (e.g., sway, balance). Methodological quality was appraised with the Physiotherapy Evidence Database (PEDro) scale and risk of bias using the Cochrane Risk of Bias 2 (RoB 2) tool. RESULTS: Fifty-one randomized controlled trials (n&#x2009;=&#x2009;2319) were included. Comparative synthesis showed that improvements inferred from surrogate outcomes were more frequent and often larger than improvements observed in direct psychophysical measures. Directly targeted practice, angle specific, attentionally demanding, and aligned with the measured proprioceptive submodality and task construct, produced the most consistent benefits in position-reproduction accuracy/error, movement-detection sensitivity, or discrimination performance, depending on the outcome assessed. In contrast, multimodal regimens (balance, strengthening, taping, manual therapy) commonly improved balance, pain, strength, or function without comparably consistent evidence of enhanced direct psychophysical proprioceptive function. CONCLUSIONS: Specific psychophysical components of proprioceptive function appear modifiable, but only when training explicitly targets the sensory construct measured. The field remains conceptually diffuse, with frequent conflation of sensorimotor performance and proprioception. Progress depends on defining proprioceptive submodalities a priori, privileging validated psychophysical outcomes over surrogate outcomes, and aligning intervention content with measurement to substantiate true perceptual learning rather than generic motor adaptation.

Journal Article

Comparison between measured and synthesized posterior lead electrocardiograms during percutaneous coronary intervention-induced myocardial ischemia.

BACKGROUND: Posterior/inferolateral myocardial ischemia is frequently underrecognized on standard 12&#x2011;lead electrocardiography (ECG). Synthesized posterior leads derived from the standard 12&#x2011;lead ECG have been proposed as an alternative to directly measured posterior leads; however, their accuracy under controlled ischemic conditions has not been fully validated. METHODS: We prospectively enrolled 26 consecutive patients undergoing percutaneous coronary intervention (PCI) in whom simultaneously recorded measured and synthesized posterior lead ECGs (V7-V9) were obtained during balloon-induced myocardial ischemia. ST-segment deviation was measured at the ST junction (STJ), 40&#xa0;ms (ST1), and 80&#xa0;ms (ST2) thereafter. Agreement between measured and synthesized posterior leads was assessed using Pearson correlation and Bland-Altman analyses. As an exploratory patient-level analysis, diagnostic performance was compared with reciprocal anterior ST-segment depression (V1-V4). RESULTS: Strong correlations were observed between measured and synthesized posterior lead ST-segment deviations (V7: r&#xa0;=&#xa0;0.89; V8: r&#xa0;=&#xa0;0.86; V9: r&#xa0;=&#xa0;0.83; all P&#xa0;<&#xa0;0.001). Bland-Altman analysis demonstrated minimal systematic bias (within &#xb1;0.004&#xa0;mV) and narrow limits of agreement. Synthesized posterior leads showed higher diagnostic performance than reciprocal anterior ST-segment depression (AUC 0.917 vs. 0.708), although the difference was not statistically significant (DeLong test, P&#xa0;=&#xa0;0.197). Using a 0.05&#xa0;mV threshold, synthesized posterior leads demonstrated 83.3% sensitivity, 100% specificity, and 96.2% overall accuracy. CONCLUSIONS: Synthesized posterior leads closely reproduced measured posterior lead ST-segment deviations during percutaneous coronary intervention (PCI)-induced myocardial ischemia, supporting the technical validity of posterior lead reconstruction. Larger prospective studies are warranted to determine whether synthesized posterior leads provide incremental diagnostic value beyond careful interpretation of the standard 12&#x2011;lead ECG.

Humans

Effectiveness of Yoga and Combined Exercise in Female With Rheumatoid Arthritis: Randomized Controlled Trial.

BACKGROUND: Although exercise is beneficial for Rheumatoid Arthritis (RA), the comparative efficacy of different modalities for patients in clinical remission remains unclear. This study compared the short- and long-term effects of yoga versus a combined exercise programme on pain, balance, mobility, fatigue, depression, and quality of life in females with RA in remission. METHODS: In this single-blind, randomized controlled trial, 74 female participants were allocated to yoga (n&#xa0;=&#xa0;25), combined exercise (n&#xa0;=&#xa0;25), or a usual care control group (n&#xa0;=&#xa0;24). The intervention groups underwent an 8-week supervised programme. Clinical assessments, including the Visual Analogue Scale (pain), Berg Balance Scale, Timed Up and Go Test, Beck Depression Inventory, Fatigue Severity Scale, and Short Form-36, were conducted at baseline, post-intervention (8&#xa0;weeks), and follow-up (20&#xa0;weeks). RESULTS: Both intervention groups demonstrated significant improvements in all outcome measures compared with the control group at post-treatment and follow-up (p&#xa0;<&#xa0;0.05). Notably, the yoga group exhibited superior outcomes compared to the combined exercise group in reducing pain intensity (median reduction of 4.00 vs. 2.00 points; p&#xa0;<&#xa0;0.001, &#x3b7;2&#xa0;=&#xa0;0.724), as well as in physical function, balance, fatigue, depression, and quality of life at the 20-week follow-up. These benefits may be partly attributed to the incorporation of breathing and relaxation techniques inherent to yoga practice. CONCLUSIONS: Both 8-week yoga and combined exercise programs are effective in managing residual symptoms in females with RA in clinical remission. However, yoga appears to provide superior benefits in pain management and psychosocial well-being, supporting its integration into multidisciplinary RA management protocols, particularly for addressing psychosocial burden in patients achieving remission. TRIAL REGISTRATION: This study was retrospectively registered at NCT07072754 (clinicaltrials.gov).

Humans

Instruments for measuring body image in breast cancer patients: a systematic review of measurement properties.

PURPOSE: To evaluate the psychometric properties of PROMs for measuring body image in breast cancer patients. METHODS: In December 2024, a psychometric systematic review was performed in the nine databases. The COSMIN checklist was employed to evaluate the methodological quality and psychometric properties of the included body image measures. The&#xa0;level of evidence was assessed using the GRADE framework, and final recommendations were formulated for the scale. RESULTS: Thirty-eight articles evaluating fifteen PROMs were included in this review. Structural validity, internal consistency, and hypothesis testing had been most frequently evaluated. Measurement error had not been assessed for all PROMs. Twelve instruments show potential application value but require further research. The BAS-BC, PSPP, and ASI-R are not recommended for use, as these instruments do not meet the strict COSMIN thresholds for full recommendation. CONCLUSION: The BIS can be recommended as a temporary screening tool for assessing body image outcome in clinical practice. The BIRS can be tentatively advised for measuring specific postoperative body image changes. However, further comprehensive studies are required to validate the psychometric properties of existing PROMs.

Female

Reliability-aware hierarchical learning for Chagas disease screening from 12-lead ECGs: tackling label uncertainty and class imbalance.

Objective.Chagas disease, a neglected tropical disease (NTD) with significant cardiovascular impact, remains underdiagnosed in resource-limited regions. Electrocardiogram (ECG) screening offers a low-cost tool for detecting cardiac involvement, yet algorithm development is challenged by label noise, data scarcity, and the latent nature of infection. This study proposes a robust ECG-based screening framework that explicitly addresses these constraints.Approach.We introduce aReliability-Aware Hierarchical Learningstrategy that calibrates supervision according to data provenance, prioritizing serology-confirmed labels over noisy self-reports. To mitigate data scarcity, we compare a specialized convolutional neural network (CNN) trained from scratch with a transfer learning approach based on a Spatio-Temporal ECG foundation Model (FM). Performance is evaluated across varying data scales, and the representation structure is analyzed to interpret model behavior.Main results.On the official hidden test set of the George B. Moody PhysioNet/Computing in Cardiology Challenge 2025, our approach achieved a Challenge Score of 0.163. We observe that while the specialized CNN performs competitively in data-rich regimes, the FM exhibits superior robustness in extreme low-resource settings. Furthermore, performance reaches a plateau imposed by underlying disease physiology. Bimodal score distributions suggest that models distinguish established cardiomyopathy from indeterminate infection, which remains electrophysiologically indistinguishable from healthy controls.Significance.These findings clarify both the potential and intrinsic limits of ECG-based AI screening for NTD-associated cardiac involvement. Reliability-aware supervision and data-efficient transfer learning provide a practical framework toward scalable and clinically meaningful ECG screening systems in resource-constrained environments.

Humans

The effects of visuomotor training and tDCS stimulation on visuomotor integration and visual processing: an electrophysiological approach.

BACKGROUND: Visuomotor integration coordinates visual and motor cortical activity to produce goal-directed responses and can be indexed by Rolandic Mu-rhythm suppression and visual evoked potential (VEP) P100 parameters. Perceptual-motor training improves visuomotor performance, and transcranial direct current stimulation (tDCS) over primary motor cortex (M1) has been reported to enhance motor learning when paired with training. This study examined whether anodal M1 tDCS augments the effects of Senaptec visuomotor training in healthy adults. METHODS: Sixty participants were randomized to active anodal tDCS (five 10-minute sessions, 1&#xa0;mA; n&#xa0;=&#xa0;31) or sham (n&#xa0;=&#xa0;29) over M1 immediately before each Senaptec training session; 53 completed all sessions and post-testing. Outcomes were Mu-suppression ratios, VEP P100 latency and amplitude, and Senaptec measures of visual sensitivity and visuomotor control. RESULTS: Active tDCS produced no augmentation of any outcome, with no significant group&#xa0;&#xd7;&#xa0;time interaction for any measure, consistent across composite and task-level analyses. Training alone produced no change in Mu suppression or visuomotor control. By contrast, both groups showed significant training-related gains in visual sensitivity, including near-far quickness and stereopsis, accompanied by shorter P100 latencies and larger amplitudes, indicating more efficient early visual processing. CONCLUSIONS: A clear dissociation emerged: training produced robust improvements in early visual processing, whereas neither tDCS nor training altered sensorimotor (Mu) or visuomotor-control measures. The tDCS results should be interpreted cautiously given the modest dose and limited power to detect small effects, rather than as evidence of inefficacy. Tablet-based perceptual training enhanced visual processing independent of neuromodulation.

Humans

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC&#x2009;=&#x2009;0.91-0.92 [0.83-0.97]; k&#x2009;=&#x2009;55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC&#x2009;=&#x2009;0.90-0.91 [0.85-0.95] (k&#x2009;=&#x2009;228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs&#x2009;=&#x2009;0.90 [0.83-0.94] (k&#x2009;=&#x2009;124) and ICC&#x2009;=&#x2009;0.91 [0.72-0.98] (k&#x2009;=&#x2009;9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load&#x2013;velocity relationship

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial.

BACKGROUND: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin. METHODS: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0-10 points), with a prespecified non-inferiority margin of -0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment. RESULTS: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was -0.335 points (95% CI, -1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of -0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group. CONCLUSIONS: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.

Humans

Patient-reported outcome measures for depression or anxiety symptoms in patients with cardiovascular disease: A COSMIN systematic review.

BACKGROUND: Depression and anxiety are common in patients with cardiovascular disease (CVD), but the measurement quality of patient-reported outcome measures (PROMs) used in this population remains unclear. This review aimed to evaluate the methodological quality, measurement properties, and certainty of evidence for depression and anxiety PROMs in adults with CVD and to inform instrument selection. METHODS: Following COSMIN and PRISMA guidance, four databases were searched from inception to February 2026. Studies assessing measurement properties of PROMs in adults with CVD were included. Methodological quality was evaluated using the COSMIN Risk of Bias checklist, and certainty of evidence was graded using an adapted GRADE approach. RESULTS: Sixty-six studies assessing 38 PROMs were included, comprising 29 generic and 9 CVD-specific instruments. Six PROMs met COSMIN Category A criteria: Cardiac Depression Scale-Short Form, Patient Health Questionnaire-9, Beck Depression Inventory-II, Hospital Anxiety and Depression Scale, Generalized Anxiety Disorder-7, and Major Depression Inventory. Four instruments were classified as Category C because of insufficient structural validity. Content-validity evidence was largely indeterminate or of limited certainty. Only 24 studies used confirmatory factor analysis or Rasch analysis, and no study assessed measurement error or responsiveness. Cross-cultural validity evidence was scarce. CONCLUSIONS: Six PROMs met Category A criteria, but selection should remain purpose- and context-specific. Particular attention should be given to somatic symptom overlap and intended clinical use. Further validation should prioritize content validity, measurement invariance, responsiveness, measurement error, and clinimetric performance.

Humans

Comparative effects of 12-week resistance training on unstable and stable surfaces on muscle stiffness, muscle co-activation, and balance in older patients with knee osteoarthritis.

OBJECTIVE: This randomized trial compared the effects of unstable resistance training (URT), involving resistance exercises on unstable surfaces, and stable resistance training (SRT), performed on stable surfaces, on muscle stiffness, co-activation, and balance in older adults with knee osteoarthritis (KOA). We hypothesized that URT would yield greater improvements by enhancing neuromuscular adaptability. METHODS: Fifty patients with KOA were randomly assigned to the URT group (n&#x202f;=&#x202f;25) or the SRT group (n&#x202f;=&#x202f;25). After attrition, 46 participants (URT: n&#x202f;=&#x202f;23; SRT: n&#x202f;=&#x202f;23) completed the intervention and were included in the final analysis. Both groups completed a 12-week supervised lower-limb resistance training program (3 sessions/week) consisting of 10 exercises performed under either unstable or stable support conditions. RESULTS: After 12 weeks of intervention, both groups showed significant reductions in pain intensity (p&#x202f;<&#x202f;0.001). However, compared with the SRT group, the URT group demonstrated significantly greater reductions in quadriceps stiffness (p&#x202f;<&#x202f;0.05), selected hamstring stiffness outcomes (p&#x202f;<&#x202f;0.05), and quadriceps-hamstring co-activation (p&#x202f;<&#x202f;0.001), alongside superior improvements in both dynamic balance and static balance (all p&#x202f;<&#x202f;0.05). CONCLUSION: While both training modalities are effective for pain relief, URT elicited greater improvements in balance-related performance and neuromuscular-mechanical outcomes than SRT in older adults with KOA. These findings suggest that incorporating unstable support conditions into resistance training may provide additional rehabilitation benefits for this population.

Humans

Comparison of black carbon measurements using filter-specific reference transmittance to those using lab blanks or an average of unloaded filters.

Filter-based optical techniques compare light transmission intensities between loaded (I) and unloaded (I0) filters as a measure of light-absorbing mass for subsequent estimations of equivalent black carbon (eBC). We analyzed 5,379 15&#x2009;mm Teflon filters from the Household Air Pollution Intervention Network (HAPIN) trial to assess the influence that different methods of I0 estimations have on eBC measures. We compared eBC measurements using filter-specific I0 values (Method 1) to those using three other methods of I0 estimation: the lab blank scan from a given session (Method 2), the average of all pre-sample filter scans (Method 3), and the average of all lab blank filter scans (Method 4). We assessed the agreement between Method 1 and the alternative methods using Bland-Altman analysis. We also assessed the relationship between Method 1 and the alternative methods across the complete measurement range and after stratifying exposure data into quartiles according to Method 1 eBC exposures. The mean (SD) personal eBC exposure for Method 1 was 7.8&#x2009;&#x3bc;g/m3 (5.9), and exposures ranged from 1.3 to 46.8&#x2009;&#x3bc;g/m3. Compared to Method 1, eBC using Methods 2, 3, and 4 were higher by 0.7&#x2009;&#x3bc;g/m3, 0.1&#x2009;&#x3bc;g/m3, and 0.7&#x2009;&#x3bc;g/m3, respectively. The performances of linear regression models between Method 1 and all other methods were moderate to strong (R2 range: 0.42-0.93) in the second, third, and fourth quartiles; however, the models in the first quartile (eBC range: 1.3-2.9&#x2009;&#x3bc;g/m3) performed poorly (R2&#x2009;=&#x2009;0.25-0.26), with error approximately 25% of the mean. Our findings suggest that, in most instances, conventional methods for obtaining I0 values can be used to sufficiently characterize eBC; however, analyzing filters before sampling adds appreciably to the accuracy of eBC estimations in lower concentration settings.Implications: Filter-based optical techniques compare light transmission intensities between loaded (I) and unloaded (I0) filters as a measure of light-absorbing mass for subsequent estimations of equivalent black carbon (eBC). To assess the influence that different methods of I0 estimations have on eBC measures, we analyzed 5,379 15 mm Teflon filters from the Household Air Pollution Intervention Network (HAPIN) trial. Our findings suggest that, in most instances, conventional methods for obtaining I0 values can be used to sufficiently characterize eBC; however, analyzing filters before sampling adds appreciably to the accuracy of eBC estimations in lower concentration settings.

Soot

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation

Clinical Utility of Ultra-Widefield Swept-Source OCT for Intraocular Tumors: Comparison With Ultrasonography, SD-OCT, and MRI.

PURPOSE: To evaluate the clinical performance of ultra-widefield swept-source optical coherence tomography (UWF-OCT) in the assessment of choroidal tumors and to compare it with ultrasonography (US), spectral-domain (SD)-OCT, and magnetic resonance imaging (MRI). DESIGN: Retrospective diagnostic comparison. SUBJECTS: Thirty-nine eyes from 39 patients diagnosed with choroidal tumors at a single tertiary referral center. METHODS: This retrospective diagnostic comparison evaluated patients diagnosed with choroidal tumors at a single tertiary referral center between January 2023 and August 2025. All patients underwent UWF-OCT imaging at diagnosis. Tumor measurements obtained with UWF-OCT were compared with US, SD-OCT, and MRI. Comparative analysis among imaging modalities and predictors affecting UWF-OCT applicability was performed. MAIN OUTCOME MEASURES: Tumor thickness (mm) and largest basal diameter (LBD, mm) measurements, and complete measurability rate across different tumor size categories. RESULTS: Thirty-nine eyes from 39 patients (mean age 59.2 &#xb1; 16.9 years) were analyzed, including 27 choroidal melanomas (69.2%), 5 metastatic tumors (12.8%), 4 hemangiomas (10.3%), 2 osteomas (5.1%), and 1 (2.6%) indeterminate choroidal melanocytic lesion. UWF-OCT successfully measured both tumor thickness and largest basal diameter (LBD) in 100% (31/31) of small and medium choroidal tumors, substantially outperforming SD-OCT (complete measurement achieved in 63.6% of small tumors, and 0% of medium or large tumors). UWF-OCT measurements were systematically smaller than ultrasonography (thickness: -32.3%, P < .01; LBD: -11.1%, P < .01) and MRI (thickness: -29.2%, P < .01). Mushroom-shaped tumor morphology was the strongest negative predictor of UWF-OCT quality (OR = 0.015, 95% CI 0.001-0.196, P < .01). UWF-OCT's complete measurability was limited in large tumors (12.5%, 1/8). CONCLUSIONS: UWF-OCT provides precise, noninvasive, single-scan assessment of small-to-medium choroidal tumors with detailed structural visualization. It may be particularly useful for dome-shaped tumors, while multimodal imaging with US and MRI remains optimal for complex morphologies. Overall, UWF-OCT represents a valuable tool for diagnosis and treatment planning, with potential utility for longitudinal follow-up in choroidal tumor management.

Humans