Academic Journal

Linguistic Fidelity and Classification Performance of Large Language Models for Generating Synthetic Operative Notes: Evaluation Study.

Bibliographic Details
Title: Linguistic Fidelity and Classification Performance of Large Language Models for Generating Synthetic Operative Notes: Evaluation Study.
Authors: Cox M; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110., Lin E; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110., Oleck N; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110., Jones C; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110., Li NY; Department of Orthopaedic Surgery, Duke University Hospital, Durham, NC, United States., Mithani SK; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110., Allori AC; Division of Plastic, Oral, and Maxillofacial Surgery, Duke University Hospital, 2301 Erwin Road, Durham, NC, 27710, United States, 1 919-668-3110.
Source: JMIR formative research [JMIR Form Res] 2026 Jul 03; Vol. 10, pp. e87276. Date of Electronic Publication: 2026 Jul 03.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: JMIR Publications Country of Publication: Canada NLM ID: 101726394 Publication Model: Electronic Cited Medium: Internet ISSN: 2561-326X (Electronic) Linking ISSN: 2561326X NLM ISO Abbreviation: JMIR Form Res Subsets: MEDLINE
Imprint Name(s): Original Publication: Toronto, ON, Canada : JMIR Publications, [2017]-
MeSH Terms: Large Language Models* , Natural Language Processing* , Linguistics*, Cleft Palate/surgery ; Cleft Lip/surgery ; Humans ; Classification Algorithms ; Machine Learning ; Reproducibility of Results
Abstract: Background: Machine learning models for surgical applications require large, diverse datasets; however, data scarcity remains a critical limitation due to privacy regulations, institutional variability, and the rarity of many surgical procedures. Large language models (LLMs) offer a potential solution through synthetic data generation, but their performance and reliability in specialized surgical domains remain underexplored.
Objective: This study aimed to evaluate the linguistic fidelity of LLM-generated operative notes for cleft lip and palate procedures and to assess their impact on natural language processing classifier performance under varying data availability conditions.
Methods: A total of 630 authentic operative notes were obtained from cleft procedures (86 primary cleft lip repairs, 101 primary cleft palate repairs, and 62 primary alveolar bone grafting [ABG] procedures) performed between 2013 and 2024. GPT-4o generated matched synthetic notes using multishot prompting with anonymized examples. Linguistic fidelity was evaluated using BERTScore for semantic similarity, Jensen-Shannon divergence of part-of-speech trigrams for syntactic structure, and Bilingual Evaluation Understudy (BLEU) scores for lexical overlap. Binary classifiers using ClinicalBERT embeddings and logistic regression were trained under both full data and data-scarce (retaining 5% or 10% of positive training cases while retaining the full negative training set) conditions, with and without synthetic augmentation at approximate ratios of synthetic to real notes (1:1, 2:1, 5:1, and 10:1).
Results: Synthetic notes demonstrated high semantic fidelity across all procedures (BERTScore F1-score: 0.86-0.88) and low syntactic divergence (Jensen-Shannon divergence: 0.06-0.08). BLEU scores indicated moderate lexical variation (0.14-0.19), reflecting distinct but contextually consistent phrasing. With full datasets, synthetic augmentation did not meaningfully affect classifier performance. Under data-scarce conditions retaining 5% of positive training cases while preserving the full negative training set, the area under the curve improved from 0.915 (SD 0.056) to 0.929 (SD 0.033) for cleft lip and from 0.935 (SD 0.036) to 0.949 (SD 0.033) for cleft palate, with smaller gains for ABG (mean 0.983, SD 0.014 to mean 0.987, SD 0.015). When retaining 10% of positive training cases, performance changes were minor across procedures (cleft lip: mean 0.952, SD 0.036 to mean 0.948, SD 0.028; cleft palate: mean 0.945, SD 0.044 to mean 0.956, SD 0.036; and ABG: mean 0.989, SD 0.008 to mean 0.985, SD 0.012).
Conclusions: LLM-generated operative notes exhibit strong semantic and syntactic fidelity to authentic documentation and can enhance model performance in a task-dependent manner when authentic data are limited. These findings suggest that synthetic data generation may address data scarcity challenges in specialized surgical domains, particularly for rare or underrepresented procedures, enabling robust machine learning model development.
(© Meredith Cox, Elaine Lin, Nicholas Oleck, Carlee Jones, Neill Y Li, Suhail K Mithani, Alexander C Allori. Originally published in JMIR Formative Research (https://formative.jmir.org).)
References: J Clin Med. 2024 May 22;13(11):. (PMID: 38892752)
JMIR Med Inform. 2020 Feb 20;8(2):e16492. (PMID: 32130148)
PLOS Digit Health. 2026 Mar 9;5(3):e0001290. (PMID: 41801961)
Annu Int Conf IEEE Eng Med Biol Soc. 2012;2012:5098-101. (PMID: 23367075)
Front Surg. 2024 May 17;11:1403540. (PMID: 38826809)
Health Policy. 2010 Oct;97(2-3):267-74. (PMID: 20800762)
J Gen Intern Med. 2019 Nov;34(11):2355-2367. (PMID: 31183688)
Health Inf Manag. 2023 Jan;52(1):18-27. (PMID: 32367733)
Neurosurg Focus. 2025 Jul 1;59(1):E17. (PMID: 40591960)
Surgery. 2024 Jun;175(6):1496-1502. (PMID: 38582732)
Clin Transl Sci. 2019 Jul;12(4):329-333. (PMID: 31074176)
NPJ Digit Med. 2020 May 14;3:69. (PMID: 32435697)
IEEE J Biomed Health Inform. 2023 Feb;27(2):790-803. (PMID: 35737624)
Nat Med. 2023 Nov;29(11):2929-2938. (PMID: 37884627)
Ann Surg. 2021 May 1;273(5):900-908. (PMID: 33074901)
Front Artif Intell. 2025 Feb 05;8:1533508. (PMID: 39974356)
BMC Public Health. 2014 Nov 05;14:1144. (PMID: 25377061)
J Med Syst. 2022 Nov 16;46(12):96. (PMID: 36380246)
Cancer Causes Control. 2019 Aug;30(8):901-908. (PMID: 31144088)
Contributed Indexing: Keywords: AI; artificial intelligence; cleft lip; cleft palate; large language models; machine learning; natural language processing
Entry Date(s): Date Created: 20260703 Date Completed: 20260703 Latest Revision: 20260813
Update Code: 20260814
PubMed Central ID: PMC13331248
DOI: 10.2196/87276
PMID: 42397952
Database: MEDLINE
Description
ISSN:2561-326X
DOI:10.2196/87276