Search PubMedSearch

PubMed · 42126183

Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.

Abstract

OBJECTIVE: Evaluate the utility of the large language model (LLM), ChatGPT, for the analysis of operative notes and the generation of Current Procedural Terminology (CPT) codes in comparison to human clinical coders. STUDY DESIGN: CPT billing codes assigned by ChatGPT were compared to existing billing data. Otology practice within a tertiary academic center. METHODS: About 191 operative notes from a single surgeon (9/2022-10/2023) were analyzed. ChatGPT-3.5 and 4 models were prompted for CPT codes based on operative notes. Assessment included determining exact and partial match rates, sensitivity and specificity for targeted procedures, and work Relative Value Units (wRVU) differences between ChatGPT-generated and human-assigned codes. RESULTS: ChatGPT-3.5 achieved exact matches in 22% of cases and partial matches in 32%, while ChatGPT-4 achieved 14% exact and 33% partial matches. When cochlear implantation (CI) was excluded, performance dropped significantly. For CI, ChatGPT-3.5 demonstrated a sensitivity of 94% and specificity of 90%, while ChatGPT-4 showed a sensitivity of 96% and specificity of 92%. In contrast, performance on cartilage grafting was poor, with sensitivities of 4.2% for ChatGPT-3.5 and 0% for ChatGPT-4. ChatGPT-3.5 and 4 showed moderate CPT code matching accuracy among themselves, with slight agreement to human coders. Both models tended to underbill for wRVUs compared to human coders, with significant differences in the values generated. CONCLUSION: This study assessed ChatGPT's effectiveness in automating CPT code assignment for otologic surgeries. While the models achieved high sensitivity values for assigning codes related to cochlear implantation, both models struggled with complex cases, failed to apply modifiers, and often assigned fewer wRVUs. The findings highlight ChatGPT's potential in medical billing but indicate a need for further refinement.

Explore related subjects

Keep this discovery

BibTeXRIS

Armo Derbarsegian, Adam S Vesole, Daniel Q Sun, Steven A Gordon. 2026-05-13. Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.. https://doi.org/10.1177/00034894261452167

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Modulating sentence comprehension in people with aphasia through anodal tDCS: A double-blind randomized cross-over study.

This double-blind randomized cross-over study investigated the effects of perilesional anodal transcranial direct current stimulation (AtDCS) combined with speech-language therapy on sentence comprehension in eight individuals with chronic nonfluent agrammatic aphasia. The behavioral therapy consisted of an intensive comprehension treatment including drilling in sentence-to-picture matching and Mapping Therapy. Each participant underwent both the anodal tDCS and sham stimulation conditions (five received sham first followed by real stimulation, and the remaining three the reverse sequence), with each condition paired with the same behavioral treatment and separated by a four-month washout period. Stimulation was applied over the perilesional area (left BA6) for 20 min during daily 40-min therapy sessions over four consecutive weeks. Sentence comprehension was assessed with the RiComprendo battery and functional communication with the Communicative Effectiveness Index (CETI). Data were analyzed using paired t-tests, Bayesian analyses, and linear mixed-effects models to control for baseline performance and individual variability. Both stimulation conditions produced significant pre-to-post improvements in sentence comprehension, particularly for syntactically complex structures such as passives and center-embedded object relatives. However, gains were overall greater following AtDCS, as reflected in larger effect sizes, stronger Bayes factors, and a significant treatment effect in the mixed-effects models. Only the AtDCS condition yielded significant improvements in self-perceived comprehension abilities on the CETI. These findings suggest that AtDCS over perilesional cortical areas may boost the effects of traditional language therapy on sentence comprehension, supporting its feasibility and potential as an adjuvant intervention in post-stroke aphasia rehabilitation.

Humans

Virtual, Augmented, and Mixed Reality Technologies in Neurosurgical Training: Enhancing Skills and Surgical Outcomes: A Systematic Review.

OBJECTIVE: To systematically review the role of virtual reality (VR), augmented reality (AR), and mixed reality (MR) in neurosurgical education and training. DESIGN: Systematic review conducted in accordance with the PRISMA guidelines. SETTING: A comprehensive search was performed across PubMed/MEDLINE, Scopus, Web of Science, and Google Scholar for English-language studies published between 1 January 2020 and 30 April 2026. PARTICIPANTS: Studies involving neurosurgeons, fellows, residents, and medical students (maximum sample size: n = 48) were included. RESULTS: Of 7,204 initially identified studies, 25 met the inclusion criteria. VR was primarily used for surgical simulation (100% of VR studies) and anatomical education (62.5%). AR demonstrated broader applications, including preoperative planning (40%) and intraoperative support (30%). MR was evenly distributed across simulation, planning, and intraoperative support (40% each). The most frequently improved outcomes were training effectiveness (52%) and technical proficiency (44%). Methodological quality scores, assessed using the Modified Medical Education Research Study Quality Instrument (MMERSQI), ranged from 39.5 to 84.5, indicating varied rigor. CONCLUSION: VR, AR, and MR technologies show potential to enhance surgical precision, technical skills, and educational outcomes in neurosurgical training. However, standardization of methodologies and cost-effective solutions remain essential. Future research should focus on long-term clinical impact and integration of AI-driven training models.

Virtual Reality