Search PubMedSearch

SEARCH · Search PubMed

Results for “editorial triage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

19 recordsLinked to original sources

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Sports Medicine

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Editorial Commentary: Stiff Patients After Rotator Cuff Repair: How Many Had Underrecognized Preoperative Adhesive Capsulitis?

Stiffness after rotator cuff repair is one of the most common sources of disability and one of the most common complications. Smoking, diabetes, Workers' compensation status, and traumatic tears are among the strongest risk factors. This is important information in setting appropriate expectations for both the patient and the surgeon preoperatively. Some of these factors are also associated with preoperative stiffness, so surgeons should maintain a high index of suspicion for concomitant adhesive capsulitis in patients presenting with limited motion, particularly in diabetic and non-English speaking populations. In cases where both a rotator cuff tear and adhesive capsulitis coexist, performing a concurrent capsular release or manipulation during the index procedure may improve functional outcomes and reduce the necessity for secondary surgical intervention.

Humans

Ancient DNA and Human Physiology.

Ancient DNA (aDNA) enables the reconstruction of chronologically sampled genomes from ancient humans, animals, plants, pathogens, and microorganisms, as well as environmental DNA, providing a record of biological changes through time. Improvements in short and degraded DNA extraction methods and low-cost sequencing now enable the generation of broad, cross-regional datasets that expand evolutionary analyses from past population demography to biological mechanisms. By tracking temporal shifts of allele frequencies, integrating functional genomics resources (e.g., gene expression, chromatin structure variation), modeling population demography to separate selection from genetic drift, and aligning genetic changes with archaeological, cultural, and climatic data, aDNA has the potential to link sequence variation to physiological function within their temporal and environmental contexts. In this review, we summarize illustrative case studies from aDNA research spanning complex traits, dietary adaptations, and responses to pathogens and other environmental changes, showing how human biology has evolved under multiple selective pressures through time. These dated signals help triage experimental work and expose mechanisms that are rare or absent in living cohorts. Although some challenges remain, such as geographic and temporal sampling disparities, limitations in data resolution and variant detection, and genotype-phenotype uncertainties, rapid methodological progress and stronger ethical frameworks are expanding what can be inferred, making aDNA a promising tool for refining physiological pathways, their timing, and their drivers.

Humans

Mendelian randomisation for rheumatology: beyond hype-what it's good for, what it can't do, and how to read it critically.

Mendelian randomisation (MR) has become abundant in the literature, with variation in quality and frequent overinterpretation of causality. This creates a problem for clinical readers, reviewers, and editors: some MR studies can sharpen causal thinking, prioritise drug targets, and challenge misleading observational claims, whereas others are little more than automated exposure-outcome scans with causal claims disproportionate to the evidence. MR can strengthen causal inference when randomised trials are impractical and conventional observational studies are vulnerable to confounding, reverse causation, or selection bias. In rheumatology, credible MR can contribute to questions about disease aetiology, modifiable risk factors, therapeutic target validation, adverse-effect anticipation, and phenotype validation. However, its interpretation depends on whether the exposure is plausibly instrumentable, whether the genetic instruments are biologically defensible, whether assumptions are interrogated in ways appropriate to the design, and whether findings are triangulated with clinical, observational, experimental, and mechanistic evidence. Instead of recapitulating all methodological issues of MR, this review aims to help rheumatologists distinguish robust MR from weak or overinterpreted analyses quickly. We provide an accessible framework for reading and triaging MR studies in rheumatology. Papers that use poorly justified instruments, treat medication use as drug-target evidence, interpret genetic liability as diagnosis, rely on mechanical sensitivity analyses, ignore prior evidence or ask no clinically meaningful question can often be passed over by readers. The goal is not to discourage MR in rheumatology, but to raise the standard; useful MR should clarify causal reasoning rather than simply generate another statistically significant association.

Journal Article

Automated CEAP Classification of Venous Duplex Reports Using Multimodal Artificial Intelligence.

OBJECTIVE: To develop and internally validate a prototype multimodal artificial intelligence system for automated CEAP (Clinical, Etiological, Anatomical and Pathophysiological) classification of venous duplex ultrasound (VDUS) reports, integrating natural language processing of free-text components with computer vision analysis of hand-drawn anatomical diagrams. METHODS: Single centre retrospective observational study using routinely collected clinical data. One thousand consecutive venous duplex ultrasound reports from Cambridge University Hospitals NHS Foundation Trust, UK (July 2024 - May 2025) were labelled according to the CEAP classification, excluding the Etiological component, which could not be reliably determined from duplex reports alone. Transfer learning was applied using ClinicalBERT for text and MobileNetV3 for diagrammatic data. Clinical classes were predicted from request line text. Text- and image-based pathophysiological models were developed for four anatomical territories (Great Saphenous Vein, Small Saphenous Vein, Deep system, Perforators), combined using late fusion with probability averaging. RESULTS: The clinical CEAP model achieved accuracy of 0.91, macro-F1 of 0.82, and macro-AUC of 0.98. Pathophysiological prediction varied, with text models broadly outperforming image models. Fusion yielded heterogeneous benefits, improving SSV performance but reducing Deep system accuracy. The performance of the final pathophysiological CEAP fusion models varied across anatomical territories: accuracy ranged from 0.70-0.92 and macro-AUC from 0.80-0.92. CONCLUSION: This study demonstrates the feasibility of automated CEAP classification from VDUS reports. Despite class imbalance affecting minority class predictions, the strong discriminatory performance validates this multimodal ML model for extracting clinically meaningful information from real-world data. This approach offers potential, pending external validation, to streamline vascular services through automated triage and guideline-compliant decision making.

Artificial intelligence

Association between packed red blood cell transfusion and clinical deterioration in neonatal necrotizing enterocolitis: a systematic review and meta-analysis.

BACKGROUND: No systematic review has evaluated the existing evidence regarding the association between packed red blood cell (pRBC) transfusion and clinical worsening of necrotizing enterocolitis (NEC) in neonates. This systematic review and meta-analysis was conducted to address this knowledge gap. MATERIALS AND METHODS: We searched the Cochrane Library, EBSCO, Embase, Web of Science, Google Scholar, and PubMed for studies on pRBC transfusion and NEC published before May 10, 2025. Relevant articles were selected through title, abstract, and full-text screening. English-language case-control studies or cohort studies, or randomized controlled trials involving newborns with NEC that compared pRBC transfusion with no transfusion and reported changes in NEC clinical status were included. Review articles, systematic reviews, case reports, editorials, animal studies, duplicate publications, and studies with incomplete data were excluded. RESULTS: Five studies involving 971 neonates with NEC were included. The pooled analysis demonstrated a potential association between pRBC transfusion and clinical deterioration of NEC in neonates (odds ratio: 6.05, 95% confidence interval: 3.02-12.14). CONCLUSIONS: pRBC transfusion was associated with an exacerbation of NEC in neonates. However, these findings should be interpreted cautiously because of the small number of eligible studies included in this meta-analysis, and future large-scale, well-designed studies are needed to confirm the observed association.

Humans

Pharmacogenomics of antipsychotic-induced weight gain: A systematic review.

BACKGROUND: Antipsychotic-induced weight gain (AIWG) is a major clinical concern, affecting approximately 30% of patients. Clinical predictors explain only part of AIWG risk. Genetic and molecular variations are hypothesized to contribute to susceptibility. The purpose of this review is to summarize recent results to identify replicated and novel findings. STUDY DESIGN: Applying PRISMA guidelines, we searched MEDLINE, Embase, and PsycINFO (May 2018-May 2026) for studies on genetic and molecular associations with AIWG, extending our prior review. Reviews, editorials, and conference abstracts were excluded. We extracted study characteristics (design, diagnosis, antipsychotic exposure, sample size, ancestry, genetic variants, and AIWG outcomes) (e.g., &#x2265;7% weight gain, BMI change). RESULTS: Fifty-three studies met inclusion criteria. In candidate gene studies, the most consistently replicated genes associated with AIWG were observed for DRD2, HTR2C, and MC4R. Multiple novel associations were identified by genome-wide association studies (GWAS) (e.g., MAP2K1, ZDBF2, PEPD), polygenic risk scores (PRS) (e.g., body mass index PRS), gene expression (e.g., CYP3A4, EP300), and epigenetic analyses (e.g., cg12034943 at CRTC1). CONCLUSIONS: Polymorphisms in candidate genes related to neurotransmission and appetite regulation continue to be investigated for associations with AIWG, while novel findings have emerged from GWAS, gene expression, and epigenetic studies. Evidence remains inconsistent due to limited replication, methodological variability, sparse ancestry data, and geographical underrepresentation. No single genetic variant is ready for clinical use, and multi-omic and multi-ancestry models are needed to improve prediction and clinical utility.

Humans

What's the meta now? More updates on the problems with systematic reviews.

BACKGROUND: Systematic reviews are intended to provide trustworthy evidence synthesis, yet previous iterations of this living review have identified numerous recurring problems in their conduct and reporting. This article presents the third version and second update of the living systematic review examining issues raised across the academic literature. METHODS: Using consistent eligibility criteria and methods from earlier versions, literature searches were updated to May 2025. Eligible meta-research and editorial articles describing problems with systematic reviews were analyzed to identify emerging themes. Additionally, four basic indicators of methodological quality of the included meta-research were presented across review versions. RESULTS: The update included 209 additional articles. Critically low methodological quality and absence of protocols remained among the most frequently reported issues in systematic reviews across disciplines and journals but notably in evidence underpinning clinical practice guidelines. Spin in abstracts and conflicts of interest continued to be common. Apparent improvements in reporting quality were inconsistent, with modest gains in some full-text reporting but persistent deficiencies in abstracts. Authorship diversity of systematic reviews improved in gender representation but remained geographically concentrated in high-income countries, and primary research included in reviews similarly lacked global representativeness. The issue of misalignment between systematic review evidence bases and global burden of disease bring the total number of problems with systematic reviews to 69. Emerging use of automation and artificial intelligence was variably reported. Descriptive comparison of meta-research articles over the three versions of this living review suggests a greater proportion meeting basic quality indicators in more recent updates. CONCLUSION: Across successive updates, problems with systematic reviews remain widespread and consistent rather than isolated. Incremental reporting improvements coexist with persistent concerns about transparency, bias, and representativeness. Future efforts should prioritize evaluating interventions and aligning research incentives to support genuinely trustworthy evidence synthesis.

Humans

Artificial intelligence-supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs.

BACKGROUND: Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers. Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways. PURPOSE: To synthesize prospective or program-embedded evaluations of AI conducted within European-style population screening programs and to estimate exploratory program-level absolute risk differences (RDs) per 1000 examinations for cancer detection rate (CDR) and recall. MATERIALS AND METHODS: We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation). Outcomes were harmonized as AI-control RDs per 1000 examinations. Random-effects pooling used Hartung-Knapp-Sidik-Jonkman models. For the paired-reader design, sensitivity analyses applied a Kish effective sample-size approach across plausible within-examination correlations (&#x3c1;&#xa0;=&#xa0;0.3-0.8). Positive predictive value (PPV) and workflow/time outcomes were summarized descriptively. RESULTS: Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I2&#xa0;&#x2248;&#xa0;12%), consistent with a modest directional increase with borderline statistical uncertainty. The pooled recall RD was -0.6 per 1000 (95% CI -3.1 to +2.1; I2&#xa0;&#x2248;&#xa0;41-43%), indicating no consistent recall increase across screening programs. Where reported, PPV was higher with AI-supported screening. Efficiency signals included 44.3% fewer total readings in MASAI and shorter reading times for AI-normal examinations in PRAIM; in PRAIM, a program-level safety-net mechanism recovered 204 cancers that would otherwise have been missed. CONCLUSION: In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (&#x2248;1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals. These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution.

Humans

Translating single-cell RNA sequencing into monocyte direct leukocyte subpopulation-transcript abundance assay ratio-based biomarkers (IFI27/PSAP or IFI27/CTSS) for clinical detection of viral infection.

A rapid method for triaging febrile patients by aetiology (e.g., viral or bacterial infection) using gene expression in peripheral blood (PB) is an intensively researched area. However, gene expression in blood represents a composite sum of gene expression of all the component cell types present in the sample. As a result, numerous genes are measured in most proposed signatures. Herein, we propose a simple ratio-based biomarker (RBB) called direct leukocyte subpopulation-transcript abundance assay (DIRECT LS-TA) that recapitulates gene expressions of a single cell type in PB (i.e., monocytes). Based on single-cell RNA sequencing (scRNAseq) data and bulk expression data, IFI27 and SIGLEC1 are found as interferon-stimulated genes (ISGs) predominantly expressed by monocytes. The DIRECT LS-TA method can use a simple ratio of two genes measured in PB as an RBB to represent the target gene expression in monocytes without the need for monocyte purification. Both scRNAseq and bulk RNA sequencing datasets were used to evaluate the correlation between ISG expression in monocytes and PB, with a particular focus on monocyte expression of IFI27. An iceberg plot of bulk transcriptome data was used to identify genes that were predominantly expressed by monocytes in PB. DIRECT LS-TA RBBs of the three genes (IFI27, IFI44L and SIGLEC1) were evaluated by group-wise comparison, receiver operating characteristic and meta-analysis. In addition, the conventional interferon (IFN) score was evaluated for comparison of diagnostic performance. In viral infection datasets, DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) was most intensely activated (p value by t test <1e-9) and had the best area under the curve (0.94) among the three potential monocyte ISGs analysed. DIRECT LS-TA SIGLEC1 was also another monocyte biomarker but showed a lower activation (p<9e-5). IFI27/PSAP showed better diagnostic performance than the conventional IFN score. On the other hand, IFI44L was not a predominant monocyte expression gene. DIRECT LS-TA of IFI27 (IFI27/PSAP or IFI27/CTSS) measured in PB was the best biomarker of viral infection and IFN activation among ISGs predominantly expressed by monocytes. It performed even better than the conventional IFN score which required quantification of eight genes. The results suggest that DIRECT LS-TA of IFI27 is a monocyte-informative biomarker which is easy to determine in PB without the need for cell sorting.

Humans

Impact of Commercial Artificial Intelligence on Radiologist Reading Time for Pulmonary Nodule Evaluation at Chest CT.

Background Chest CT is a primary method for identifying pulmonary nodules, yet interpreting scans remains time-intensive and demanding. Currently, artificial intelligence (AI) is expected to reduce reading times, but the effect of AI on reporting times in this setting is unknown. Purpose To evaluate the impact of a commercial AI software on radiologists' reading time for pulmonary nodule assessment on chest CT scans within a real-world clinical setting. Materials and Methods This retrospective study included patients who underwent chest CT examinations at a tertiary medical center between September 2021 and May 2024. The study period was divided into pre- and post-AI phases. The primary outcome was radiology reporting time. The association between AI implementation and reporting time was evaluated using a multivariable parametric Weibull shared frailty survival model adjusted for reader function, examination type, patient location, and requesting specialty, with clustering at the radiologist level. Interaction analyses assessed heterogeneity across prespecified subgroups. An exploratory extrapolation estimated projected workforce and financial impact. Results This study included 19&#x2009;433 patients (mean age, 62 years &#xb1; 14.2 [SD]; 21&#x2009;814 men; 39&#x2009;323 chest CT examinations, 19&#x2009;190 pre-AI, and 20&#x2009;133 post-AI). AI implementation was associated with faster report completion (adjusted hazard ratio, 1.17; 95% CI: 1.14, 1.21; P < .001). The adjusted median reporting time decreased from 21.3 minutes pre-AI to 18.2 minutes post-AI (14.6% reduction; P < .001). Heterogeneity was observed across reader function (P < .001), examination type (P = .048), and requesting specialty (P = .03). The largest relative reductions were observed for CT thorax electrocardiogram-gated examinations (-41.1%; P < .001) and thoracic radiologists (-25.0%; P < .001), whereas emergency department examinations showed increased median reporting time (7.1%; P < .001). At institutional scan volumes (approximately 20&#x2009;000-22&#x2009;000 chest CT examinations annually), exploratory modeling suggested an approximate reduction of 0.5 full-time equivalent radiologist workload. Conclusion Implementation of commercial AI-assisted pulmonary nodule assessment on chest CT scans reduced radiologist reporting time in a real-world clinical setting. &#xa9; The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article. See also the editorial by Iwasawa in this issue.

Humans

Effect of a digitally augmented general health promotion intervention on abstinence from health-risk behaviors among emergency department discharge patients: A randomized controlled trial.

BACKGROUND: Noncommunicable diseases (NCDs) are the leading global cause of death and are driven by modifiable behaviors, such as tobacco use, harmful alcohol consumption, unhealthy diet, and physical inactivity. Recognizing that emergency department (ED) visits represent a unique opportunity to promote behavior change, this trial evaluated a digitally augmented, theory based general health promotion approach, combining a brief telephone-based intervention with mobile instant messaging support, to help discharged ED patients abstain from health risk behaviors. METHODS AND FINDINGS: This assessor-blinded randomized controlled trial was conducted in a major public hospital ED in Hong Kong. Adults (18-65 years) triaged as semi-urgent or non-urgent and with &#x2265;1 health-risk behavior and smartphone access were randomized to receive a digitally augmented, theory&#x2011;based general health&#x2011;promotion intervention consisting of a brief telephone&#x2011;based AWARD&#x2011;model intervention (Ask, Warn, Advise, Refer, and Do-it-again) followed by weekly WhatsApp or WeChat messages for 6 months, or to a control group receiving brief telephone advice only. The primary outcome was self-report abstinence from &#x2265;1 health-risk behavior at 6 months; secondary outcomes included the proportion of participants who achieved self-reported abstinence from &#x2265;1 health-risk behavior at 12 months and reduction in the number of behaviors at 6 and 12 months. Of the 2,134 screened patients, 572 were enrolled (286 per group). At 6 months, 30.1% of the intervention participants versus 19.9% of the controls achieved self-reported abstinence (RR&#x2009;=&#x2009;1.51; 95% CI, 1.13-2.02; P&#x2009;=&#x2009;0.006). The intervention also significantly increased the likelihood of fewer risky behaviors at 6 (RR&#x2009;=&#x2009;1.54; P&#x2009;=&#x2009;0.01) and 12 (RR&#x2009;=&#x2009;1.48; P&#x2009;=&#x2009;0.02) months. Physical inactivity showed the greatest improvement at 6 months (31.7% versus 16.2%; P&#x2009;<&#x2009;0.001). The effects attenuated after cessation of booster messaging. Limitations include reliance on self-reported outcomes, the single-center study design, and loss to follow-up, which may have affected the generalizability of the results. CONCLUSIONS: A digitally augmented, theory-based general health promotion strategy delivered at ED discharge through brief telephone intervention and mobile instant messaging support demonstrated short-term benefits in promoting self-reported abstinence and reducing health-risk behaviors at 6 months. However, the absence of a sustained effect at 12 months suggests that extended support or maintenance strategies may be required to maintain these improvements over time. Multicenter trials with longer follow-up are warranted to evaluate long-term effectiveness. CLINICAL TRIAL REGISTRATION: ClinicalTrials.gov (Registration No: NCT06077565).

Humans