SegORG: Report Generation of Oral Potentially Malignant Disorders Image Based on Lesion Segmentation-Enhanced LLM.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: SegORG: Report Generation of Oral Potentially Malignant Disorders Image Based on Lesion Segmentation-Enhanced LLM.
Συγγραφείς: Zhang R; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China.; Life Health Innovation and Entrepreneurship Center, Institute of Wenzhou, Zhejiang University, Wenzhou, China.; Zhejiang Provincial Key Laboratory of Internet Multimedia Technology, College of Biomedical Engineering & Instrument Science, Zhejiang University, Hangzhou, China., Huang P; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China., Ding T; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China., Chen Y; Zhejiang Provincial Key Laboratory of Internet Multimedia Technology, College of Biomedical Engineering & Instrument Science, Zhejiang University, Hangzhou, China., Tian X; Zhejiang Provincial Key Laboratory of Internet Multimedia Technology, College of Biomedical Engineering & Instrument Science, Zhejiang University, Hangzhou, China., Cao Y; State Key Laboratory of Industrial Control Technology, College of Control Science and Engineering, Zhejiang University, Hangzhou, China., Chen W; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China., Chen X; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China., Chen Q; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China., Zhu F; Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China.
Πηγή: Oral diseases [Oral Dis] 2026 Mar; Vol. 32 (3), pp. 760-771. Date of Electronic Publication: 2025 Nov 28.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Munksgaard Country of Publication: Denmark NLM ID: 9508565 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1601-0825 (Electronic) Linking ISSN: 1354523X NLM ISO Abbreviation: Oral Dis Subsets: MEDLINE
Imprint Name(s): Publication: 2001- : Copenhagen, Denmark : Munksgaard
Original Publication: Houndmills, Basingstoke, Hampshire, UK : Stockton Press, c1995-
Ιατρικοί όροι (MeSH): Mouth Neoplasms*/diagnostic imaging , Mouth Neoplasms*/pathology , Precancerous Conditions*/diagnostic imaging , Image Processing, Computer-Assisted*/methods , Natural Language Processing*, Mouth Mucosa/diagnostic imaging ; Mouth Mucosa/pathology ; Humans
Περίληψη: Objectives: To develop an automated system for generating standardized reports for oral potentially malignant disorders (OPMDs) from white-light images, aiming to reduce documentation workload, facilitate early intervention, and enable longitudinal lesion monitoring.
Methods: We proposed the SegORG model using 441 oral mucosa images and 1323 corresponding reports. It employed SegFormer for lesion segmentation; a visual encoder extracted global and local visual embeddings, which were projected into a pre-trained large language model (LLM feature) space via a lightweight visual mapper. The Qwen2.5-7B model then generated structured diagnostic reports, enhanced by text augmentation techniques to improve diversity and professionalism.
Results: SegORG achieved BLEU-4, ROUGE-L, and CIDEr scores of 0.291, 0.517, and 0.578, respectively. Additionally, the model obtained a clinical diagnostic F1-score of 0.695 and a median expert rating of 4 on the Likert scale (p < 0.001). It significantly outperformed conventional baseline models (R2Gen, METransformer, and SwinB+BERT9k) and contemporary general-purpose multimodal LLMs (GPT-4, Gemini 2.5 Pro, and Qwen2.5-VL), despite its lean architecture (90.4 M trainable parameters).
Conclusions: By enhancing visual feature extraction and achieving efficient text alignment, SegORG offers an effective technical pathway for OPMDs reports automation. While single-center validation shows promise, multicenter trials are needed to assess generalizability.
(© 2025 John Wiley & Sons Ltd.)
References: Aksoy, N., N. Ravikumar, and A. F. Frangi. 2023. “Radiology Report Generation Using Transformers Conditioned With Non‐Imaging Data. Paper Presented at the Medical Imaging 2023: Imaging Informatics for Healthcare, Research, and Applications.”.
Chen, Z., Y. Shen, Y. Song, and X. Wan. 2022. “Cross‐Modal Memory Networks for Radiology Report Generation. arXiv.” https://ui.adsabs.harvard.edu/abs/2022arXiv220413258C. https://doi.org/10.48550/arXiv.2204.13258.
Chen, Z., Y. Song, T.‐H. Chang, and X. Wan. 2020. “Generating Radiology Reports via Memory‐Driven Transformer. Paper Presented at the Conference on Empirical Methods in Natural Language Processing (EMNLP).”.
Dashti, M., S. Ghasemi, N. Ghadimi, et al. 2024. “Performance of ChatGPT 3.5 and 4 on U.S. Dental Examinations: The INBDE, ADAT, and DAT.” Imaging Science in Dentistry 54, no. 3: 271–275. https://doi.org/10.5624/isd.20240037.
DeepSeek‐AI. 2024. “DeepSeek‐V3 Technical Report. arXiv.” https://ui.adsabs.harvard.edu/abs/2024arXiv241219437D. https://doi.org/10.48550/arXiv.2412.19437.
Deng, J., W. Dong, R. Socher, L. J. Li, L. Kai, and F.‐F. Li. 2009. “ImageNet: A Large‐Scale Hierarchical Image Database. Paper Presented at the 2009 IEEE Conference on Computer Vision and Pattern Recognition.”.
Fu, Q., Y. Chen, Z. Li, et al. 2020. “A Deep Learning Algorithm for Detection of Oral Cavity Squamous Cell Carcinoma From Photographic Images: A Retrospective Study.” EClinicalMedicine 27: 100558. https://doi.org/10.1016/j.eclinm.2020.100558.
Gallifant, J., M. Afshar, S. Ameen, et al. 2025. “The TRIPOD‐LLM Reporting Guideline for Studies Using Large Language Models.” Nature Medicine 31, no. 1: 60–69. https://doi.org/10.1038/s41591‐024‐03425‐5.
Ganesan, K. 2018. “ROUGE 2.0: Updated and Improved Measures for Evaluation of Summarization Tasks. arXiv.” https://ui.adsabs.harvard.edu/abs/2018arXiv180301937G. https://doi.org/10.48550/arXiv.1803.01937.
Guo, L., A. M. Tahir, D. Zhang, Z. J. Wang, and R. K. Ward. 2024. “Automatic Medical Report Generation: Methods and Applications.” APSIPA Transactions on Signal and Information Processing 13, no. 1: e24. https://doi.org/10.1561/116.20240044.
He, K., C. Gan, Z. Li, et al. 2023. “Transformers in Medical Image Analysis.” Intelligent Medicine 3, no. 1: 59–78. https://doi.org/10.1016/j.imed.2022.07.002.
Hou, X., Z. Liu, X. Li, X. Li, S. Sang, and Y. Zhang. 2023. “MKCL: Medical Knowledge With Contrastive Learning Model for Radiology Report Generation.” Journal of Biomedical Informatics 146: 104496. https://doi.org/10.1016/j.jbi.2023.104496.
Irvin, J., P. Rajpurkar, M. Ko, et al. 2019. “Chexpert: A Large Chest Radiograph Dataset With Uncertainty Labels and Expert Comparison. Paper Presented at the Proceedings of the AAAI Conference on Artificial Intelligence.”.
Jin, H., H. Che, Y. Lin, and H. Chen. 2023. “PromptMRG: Diagnosis‐Driven Prompts for Medical Report Generation. arXiv.” https://ui.adsabs.harvard.edu/abs/2023arXiv230812604J. https://doi.org/10.48550/arXiv.2308.12604.
Jing, B., P. Xie, and E. Xing. 2017. “On the Automatic Generation of Medical Imaging Reports. arXiv.” https://ui.adsabs.harvard.edu/abs/2017arXiv171108195J. https://doi.org/10.48550/arXiv.1711.08195.
Jubair, F., O. Al‐karadsheh, D. Malamos, S. Al Mahdi, Y. Saad, and Y. Hassona. 2022. “A Novel Lightweight Deep Convolutional Neural Network for Early Detection of Oral Cancer.” Oral Diseases 28, no. 4: 1123–1130. https://doi.org/10.1111/odi.13825.
Kirillov, A., E. Mintun, N. Ravi, et al. 2023. “Segment Anything. Paper Presented at the 2023 IEEE/CVF International Conference on Computer Vision (ICCV).”.
Kojima, T., S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa. 2022. “Large Language Models Are Zero‐Shot Reasoners. arXiv.” https://ui.adsabs.harvard.edu/abs/2022arXiv220511916K. https://doi.org/10.48550/arXiv.2205.11916.
Li, J., W. Y. Kot, C. P. McGrath, B. W. A. Chan, J. W. K. Ho, and L. W. Zheng. 2024. “Diagnostic Accuracy of Artificial Intelligence Assisted Clinical Imaging in the Detection of Oral Potentially Malignant Disorders and Oral Cancer: A Systematic Review and Meta‐Analysis.” International Journal of Surgery 110, no. 8: 5034–5046. https://doi.org/10.1097/JS9.0000000000001469.
Li, J., D. Li, S. Savarese, and S. Hoi. 2023. “BLIP‐2: Bootstrapping Language‐Image Pre‐Training With Frozen Image Encoders and Large Language Models. arXiv.” https://ui.adsabs.harvard.edu/abs/2023arXiv230112597L. https://doi.org/10.48550/arXiv.2301.12597.
Li, J., T. Su, B. Zhao, et al. 2024. “Ultrasound Report Generation With Cross‐Modality Feature Alignment via Unsupervised Guidance. arXiv.” https://ui.adsabs.harvard.edu/abs/2024arXiv240600644L. https://doi.org/10.48550/arXiv.2406.00644.
Li, M., R. Liu, F. Wang, X. Chang, and X. Liang. 2023. “Auxiliary Signal‐Guided Knowledge Encoder‐Decoder for Medical Report Generation.” World Wide Web 26, no. 1: 253–270. https://doi.org/10.1007/s11280‐022‐01013‐6.
Lingen, M. W., E. Abt, N. Agrawal, et al. 2017. “Evidence‐Based Clinical Practice Guideline for the Evaluation of Potentially Malignant Disorders in the Oral Cavity: A Report of the American Dental Association.” Journal of the American Dental Association 148, no. 10: 712–727. https://doi.org/10.1016/j.adaj.2017.07.032.
Liu, Z., Y. Lin, Y. Cao, et al. 2021. “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. arXiv.” https://ui.adsabs.harvard.edu/abs/2021arXiv210314030L. https://doi.org/10.48550/arXiv.2103.14030.
Loshchilov, I., and F. Hutter. 2017. “Decoupled Weight Decay Regularization. arXiv.” https://ui.adsabs.harvard.edu/abs/2017arXiv171105101L. https://doi.org/10.48550/arXiv.1711.05101.
Ma, J., Y. He, F. Li, L. Han, C. You, and B. Wang. 2024. “Segment Anything in Medical Images.” Nature Communications 15, no. 1: 654. https://doi.org/10.1038/s41467‐024‐44824‐z.
OpenAI. 2023. “GPT‐4 Technical Report. arXiv.” https://ui.adsabs.harvard.edu/abs/2023arXiv230308774O. https://doi.org/10.48550/arXiv.2303.08774.
Pang, T., P. Li, and L. Zhao. 2023. “A Survey on Automatic Generation of Medical Imaging Reports Based on Deep Learning.” Biomedical Engineering Online 22, no. 1: 48. https://doi.org/10.1186/s12938‐023‐01113‐y.
Papineni, K., S. Roukos, T. Ward, and W.‐J. Zhu. 2002. “Bleu: A Method for Automatic Evaluation of Machine Translation. Paper Presented at the Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, Pennsylvania, USA.”.
Parak, U., A. Lopes Carvalho, F. Roitberg, and O. Mandrik. 2022. “Effectiveness of Screening for Oral Cancer and Oral Potentially Malignant Disorders (OPMD): A Systematic Review.” Preventive Medical Reports 30: 101987. https://doi.org/10.1016/j.pmedr.2022.101987.
Parres, D., A. Albiol, and R. Paredes. 2024. “Improving Radiology Report Generation Quality and Diversity Through Reinforcement Learning and Text Augmentation.” Bioengineering 11, no. 4: 351. https://doi.org/10.3390/bioengineering11040351.
Qwen. 2024. “Qwen2.5 Technical Report. arXiv.” https://ui.adsabs.harvard.edu/abs/2024arXiv241215115Q. https://doi.org/10.48550/arXiv.2412.15115.
Ranjit, M., G. Ganapathy, R. Manuel, and T. J. Ganu. 2023. “Retrieval Augmented Chest X‐Ray Report Generation Using OpenAI GPT Models. arXiv.” https://ui.adsabs.harvard.edu/abs/2023arXiv230503660R. https://doi.org/10.48550/arXiv.2305.03660.
Schaffer, J., J. O'Donovan, J. Michaelis, A. Raglin, and T. Höllerer. 2019. “I Can Do Better Than Your AI: Expertise and Explanations. Paper Presented at the Proceedings of the 24th International Conference on Intelligent User Interfaces. Marina del Ray, California.” https://doi.org/10.1145/3301275.3302308.
Sloan, P., P. Clatworthy, E. Simpson, and M. Mirmehdi. 2025. “Automated Radiology Report Generation: A Review of Recent Advances.” IEEE Reviews in Biomedical Engineering 18: 368–387. https://doi.org/10.1109/RBME.2024.3408456.
Team Gemini. 2023. “Gemini: A Family of Highly Capable Multimodal Models. arXiv.” https://ui.adsabs.harvard.edu/abs/2023arXiv231211805G. https://doi.org/10.48550/arXiv.2312.11805.
Vedantam, R., C. L. Zitnick, and D. Parikh. 2015. “CIDEr: Consensus‐Based Image Description Evaluation. Paper Presented at the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).”.
Wang, J., A. Bhalerao, T. Yin, S. See, and Y. He. 2024. “CAMANet: Class Activation Map Guided Attention Network for Radiology Report Generation.” IEEE Journal of Biomedical and Health Informatics 28, no. 4: 2199–2210. https://doi.org/10.1109/JBHI.2024.3354712.
Wang, P., S. Bai, S. Tan, S. Wang, Z. Fan, and J. Bai. 2024. “Qwen2‐VL: Enhancing Vision‐Language Model's Perception of the World at Any Resolution. arXiv.” https://ui.adsabs.harvard.edu/abs/2024arXiv240912191W. https://doi.org/10.48550/arXiv.2409.12191.
Wang, S., B. Peng, Y. Liu, and Q. Peng. 2023. “Fine‐Grained Medical Vision‐Language Representation Learning for Radiology Report Generation. Paper Presented at the Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore.”.
Wang, Z., L. Liu, L. Wang, and L. Zhou. 2023a. “METransformer: Radiology Report Generation by Transformer With Multiple Learnable Expert Tokens. Paper Presented at the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).”.
Wang, Z., L. Liu, L. Wang, and L. Zhou. 2023b. “R2GenGPT: Radiology Report Generation With Frozen LLMs.” Meta‐Radiology 1, no. 3: 100033. https://doi.org/10.1016/j.metrad.2023.100033.
Warin, K., W. Limprasert, S. Suebnukarn, S. Jinaporntham, and P. Jantana. 2021. “Automatic Classification and Detection of Oral Cancer in Photographic Images Using Deep Learning Algorithms.” Journal of Oral Pathology & Medicine 50, no. 9: 911–918. https://doi.org/10.1111/jop.13227.
Warin, K., W. Limprasert, S. Suebnukarn, S. Jinaporntham, and P. Jantana. 2022. “Performance of Deep Convolutional Neural Network for Classification and Detection of Oral Potentially Malignant Disorders in Photographic Images.” International Journal of Oral and Maxillofacial Surgery 51, no. 5: 699–704. https://doi.org/10.1016/j.ijom.2021.09.001.
You, K., J. Gu, J. Ham, et al. 2023. “CXR‐CLIP: Toward Large Scale Chest X‐Ray Language‐Image Pre‐Training. Paper Presented at the Medical Image Computing and Computer Assisted Intervention – MICCAI 2023, Cham.”.
Zhan, Z. Z., Y. T. Xiong, C. Y. Wang, et al. 2025. “Utilizing GPT‐4 to Interpret Oral Mucosal Disease Photographs for Structured Report Generation.” Scientific Reports 15, no. 1: 5187. https://doi.org/10.1038/s41598‐025‐89328‐y.
Zhang, R., M. Lu, J. Zhang, et al. 2024. “Research and Application of Deep Learning Models With Multi‐Scale Feature Fusion for Lesion Segmentation in Oral Mucosal Diseases.” Bioengineering 11, no. 11: 1107. https://doi.org/10.3390/bioengineering11111107.
Zhang, Z., B. Wang, W. Liang, et al. 2024. “Sam‐Guided Enhanced Fine‐Grained Encoding With Mixed Semantic Learning for Medical Image Captioning. Paper Presented at the ICASSP 2024–2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).”.
Zhao, R., X. Wang, H. Dai, P. Gao, and P. Li. 2023. “Medical Report Generation Based on Segment‐Enhanced Contrastive Representation Learning. Paper Presented at the Natural Language Processing and Chinese Computing, Cham.”.
Zhou, Y., A. Ioan Muresanu, Z. Han, et al. 2022. “Large Language Models Are Human‐Level Prompt Engineers. arXiv.” https://ui.adsabs.harvard.edu/abs/2022arXiv221101910Z. https://doi.org/10.48550/arXiv.2211.01910.
Grant Information: 2024YFC2510700 the National Key R&D Program of China; XMGL-KJZX-202401 Science and Technology Special Project, Institute of Wenzhou, Zhejiang University; 82471034 the National Natural Science Foundation of China; 82301070 the National Natural Science Foundation of China
Contributed Indexing: Keywords: large language models; lesion segmentation; medical report generation; oral potentially malignant disorders; text augmentation; vision transformer
Entry Date(s): Date Created: 20251128 Date Completed: 20260429 Latest Revision: 20260429
Update Code: 20260429
DOI: 10.1111/odi.70148
PMID: 41313616
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1601-0825
DOI:10.1111/odi.70148