Academic Journal
Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models: Comparative Study.
| Τίτλος: | Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models: Comparative Study. |
|---|---|
| Συγγραφείς: | Ahmed S; Department of Computer Science & Engineering, BRAC University, Kha 224 Pragati Sarani, Merul Badda, Dhaka, 1212, Bangladesh, 880 1796965173., Yousuf Sadeque F; Department of Computer Science & Engineering, BRAC University, Kha 224 Pragati Sarani, Merul Badda, Dhaka, 1212, Bangladesh, 880 1796965173. |
| Πηγή: | JMIR medical informatics [JMIR Med Inform] 2026 Jun 03; Vol. 14, pp. e82545. Date of Electronic Publication: 2026 Jun 03. |
| Τύπος έκδοσης: | Journal Article; Comparative Study |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: JMIR Publications Country of Publication: Canada NLM ID: 101645109 Publication Model: Electronic Cited Medium: Internet ISSN: 2291-9694 (Electronic) Linking ISSN: 22919694 NLM ISO Abbreviation: JMIR Med Inform Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: Toronto : JMIR Publications, [2013]- |
| Ιατρικοί όροι (MeSH): | Large Language Models* , Physician-Patient Relations* , Natural Language Processing* , Communication*, Humans |
| Περίληψη: | Background: Clinical note documentation is a vital yet time-intensive task in health care. While advancements in natural language processing have transformed many domains, generating accurate summaries of doctor-patient conversations remains underexplored due to the limited availability of open-source datasets. Large language models (LLMs), with their training on vast datasets, present a promising solution to this challenge. Objective: Precision in clinical summarization is crucial, as it directly impacts patient care and safety. This study aimed to evaluate the effectiveness of parameter-efficient, fine-tuned, decoder-only LLMs for clinical note generation from doctor-patient conversations. We focus on assessing medical accuracy, robustness, and the feasibility of parameter-efficient fine-tuning (PEFT) approaches under practical resource constraints. Methods: We used the Medical Training Summarization Dialog dataset containing 1700 doctor-patient conversations paired with clinical notes. Several decoder-only LLMs, including Mistral, Meditron, and Llama, were fine-tuned using PEFT techniques to reduce computational and memory overhead. Evaluation was performed using standard automatic metrics, including the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers score, to assess content overlap and semantic similarity between generated and reference clinical notes. In addition, an expert physician assessed the LLM-generated notes for medical accuracy, completeness, concision, relevance, and clinical coherence and readability. Results: Model performance was evaluated using the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers scores, demonstrating that Meditron-7B and Llama3-8B achieved state-of-the-art results among open-source, parameter-efficient, fine-tuned models, with Mistral-7B also performing competitively. The findings indicate that decoder-only LLMs, particularly Llama variants, outperform traditional models. Moreover, fine-tuning with higher quantization has the potential to further enhance performance. Human expert evaluation further indicated that Llama3-8B and Mistral-7B produced clinically coherent and accurate summaries, with Meditron-7B and Llama3-3B also performing reliably across evaluation criteria. The findings suggest that higher quantization during fine-tuning may improve efficiency without substantially compromising performance. Conclusions: This study underscores the potential of the PEFT of decoder-only LLMs to transform clinical workflows by streamlining medical documentation, thereby enabling health care professionals to dedicate more time to patient care. These models offer a scalable and resource-efficient alternative to traditional architectures and have the potential to streamline clinical documentation workflows. (© Saib Ahmed, Farig Yousuf Sadeque. Originally published in JMIR Medical Informatics (https://medinform.jmir.org).) |
| References: | J Am Med Inform Assoc. 2011 Mar-Apr;18(2):112-7. (PMID: 21292706) J Am Med Inform Assoc. 2020 Jul 1;27(7):1132-1135. (PMID: 32324855) JMIR Med Inform. 2024 Aug 28;12:e59617. (PMID: 39195570) J Med Internet Res. 2025 Sep 23;27:e76048. (PMID: 40986888) |
| Contributed Indexing: | Keywords: BERTScore; Dialogue2Note; Llama; Meditron; Mistral; ROUGE score; Recall-Oriented Understudy for Gisting Evaluation; Recall-Oriented Understudy for Gisting Evaluation score; bidirectional encoder representations from transformers; bidirectional encoder representations from transformers score; clinical NLP; clinical natural language processing; decoder-only; natural language processing; summarization; transformer |
| Entry Date(s): | Date Created: 20260603 Date Completed: 20260611 Latest Revision: 20260813 |
| Update Code: | 20260813 |
| PubMed Central ID: | PMC13232911 |
| DOI: | 10.2196/82545 |
| PMID: | 42234930 |
| Βάση Δεδομένων: | MEDLINE |
| ISSN: | 2291-9694 |
|---|---|
| DOI: | 10.2196/82545 |