Search PubMedSearch

PubMed · 42438216

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

Abstract

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Explore related subjects

Keep this discovery

BibTeXRIS

Romir Patel, Christopher Shultz, Matthieu Ollivier, Daniel C Wascher. 2026-07-12. Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.. https://doi.org/10.1177/03635465261463006

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Self-Report Health Screening Tools in Female Athletes: A Systematic Review of Domain Coverage, Validation, and Use Across Participation Levels.

BACKGROUND: Female athlete health encompasses multiple interconnected domains; however, the self-report screening tools used to assess these domains have not been comprehensively synthesised. OBJECTIVE: To systematically identify self-report health screening tools used to assess female athlete health, map domain coverage, determine validation reporting, and describe application across participation levels. METHODS: This systematic review was pre-registered with PROSPERO ( CRD420251056910 ) and conducted in accordance with PRISMA guidelines. Four databases (PubMed, MEDLINE, SPORTDiscus and Web of Science) were searched from inception to January 2026 using female health and screening-related terms. Methodological quality was appraised using Joanna Briggs Institute and National Institutes of Health tools, and findings were synthesised descriptively. Eligible, peer-reviewed studies reported the use, development or validation of self-report health screening tools assessing one or more domains relevant to female health applied in female athlete populations, spanning recreational through elite participation levels. All sports and activities were included. The search was restricted to English language with no date limits. RESULTS: In total, 360 studies (1990-2026) representing 134,506 female participants spanning recreational to elite sport and 273 screening tools were included. Mental health (n&#x2009;=&#x2009;77, 34.1%), disordered eating (n&#x2009;=&#x2009;33, 14.6%) and body image (n&#x2009;=&#x2009;30, 13.3%) predominated. Domains related to female health, including menstrual health, pelvic floor health, pregnancy/postpartum and breast health were comparatively underrepresented. Most&#xa0;studies reported tools were used for risk identification (n&#x2009;=&#x2009;323,&#xa0;80.3%). Validation reporting was inconsistent, with half (n&#x2009;=&#x2009;180,&#xa0;50%) reporting use of at least one validated tool. Tool use was concentrated in professional and elite sport, with limited inclusion of recreational, masters and disability athlete cohorts. Health literacy constructs were explicitly&#xa0;assessed in 12.5% of studies&#xa0;(n&#x2009;=&#x2009;45). CONCLUSIONS: Health screening in female athlete populations remains fragmented and uneven in domain coverage, with inconsistent validation reporting. Development of integrated, multi-domain and contextually inclusive screening frameworks is warranted.

Journal Article

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans