Academic Journal

Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy.
Συγγραφείς: Liu Y; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China.; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Li R; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China.; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Yu P; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China.; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Zhang Y; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China.; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Zhao A; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Xiao X; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China., Li Z; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China., Liang R; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China. 18326900957@163.com., Peng L; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China. penglei933@163.com.; Department of Urology, Medical School, South China Hospital, Shenzhen University, Shenzhen, Guangdong, China. penglei933@163.com., Dong Z; Department of Urology, The Second Hospital & Clinical Medical School, Lanzhou University, Gansu, China. dzl19780829@163.com.
Πηγή: Journal of robotic surgery [J Robot Surg] 2026 Jul 22; Vol. 20 (1). Date of Electronic Publication: 2026 Jul 22.
Τύπος έκδοσης: Journal Article; Comparative Study
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Springer Country of Publication: England NLM ID: 101300401 Publication Model: Electronic Cited Medium: Internet ISSN: 1863-2491 (Electronic) Linking ISSN: 18632483 NLM ISO Abbreviation: J Robot Surg Subsets: MEDLINE
Imprint Name(s): Original Publication: London : Springer
Ιατρικοί όροι (MeSH): Patient Education as Topic*/methods , Robotic Surgical Procedures*/methods , Cystectomy*/methods , Comprehension* , Artificial Intelligence*, Humans ; Cross-Sectional Studies ; Reproducibility of Results
Περίληψη: Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.
(© 2026. The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature.)
Competing Interests: Declarations. Ethics approval and consent to participate: Not applicable. Competing interests: The authors declare no competing interests.
References: Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I et al (2024) Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. Cancer J Clin 74(3):229–263. https://doi.org/10.3322/caac.21834. (PMID: 10.3322/caac.21834)
Powles T, Bellmunt J, Comperat E, De Santis M, Huddart R, Loriot Y et al (2022) Bladder cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Annals oncology: official J Eur Soc Med Oncol 33(3):244–258. https://doi.org/10.1016/j.annonc.2021.11.012. (PMID: 10.1016/j.annonc.2021.11.012)
van der Heijden AG, Bruins HM, Carrion A, Cathomas R, Compérat E, Dimitropoulos K et al (2025) European Association of Urology Guidelines on Muscle-invasive and Metastatic Bladder Cancer: Summary of the 2025 Guidelines. Eur Urol 87(5):582–600. https://doi.org/10.1016/j.eururo.2025.02.019. (PMID: 10.1016/j.eururo.2025.02.01940118736)
Siracusano S, Gontero P, Mearini E, Imbimbo C, Simonato A, Dal Moro F et al (2024) Short-term effects of bowel function on global health quality of life after radical cystectomy. Minerva Urol Nephrol 76(4):452–457. https://doi.org/10.23736/s2724-6051.24.05730-6. (PMID: 10.23736/s2724-6051.24.05730-638842052)
Parekh DJ, Reis IM, Castle EP, Gonzalgo ML, Woods ME, Svatek RS et al (2018) Robot-assisted radical cystectomy versus open radical cystectomy in patients with bladder cancer (RAZOR): an open-label, randomised, phase 3, non-inferiority trial. Lancet (London England) 391(10139):2525–2536. https://doi.org/10.1016/s0140-6736(18)30996-6. (PMID: 10.1016/s0140-6736(18)30996-629976469)
Catto JWF, Khetrapal P, Ricciardi F, Ambler G, Williams NR, Al-Hammouri T et al (2022) Effect of robot-assisted radical cystectomy with intracorporeal urinary diversion vs open radical cystectomy on 90-day morbidity and mortality among patients with bladder cancer: a randomized clinical trial. JAMA 327(21):2092–2103. https://doi.org/10.1001/jama.2022.7393. (PMID: 10.1001/jama.2022.7393355690799109000)
Bochner BH, Dalbagni G, Sjoberg DD, Silberstein J, Keren Paz GE, Donat SM et al (2015) Comparing open radical cystectomy and robot-assisted laparoscopic radical cystectomy: a randomized clinical trial. Eur Urol 67(6):1042–1050. https://doi.org/10.1016/j.eururo.2014.11.043. (PMID: 10.1016/j.eururo.2014.11.04325496767)
Pandolfo SD, Aveta A, Loizzo D, Crocerossa F, La Rocca R, Del Giudice F et al (2022) Quality of web-based patient information on robotic radical cystectomy remains poor: a standardized assessment. Urol Pract 9(5):498–503. https://doi.org/10.1097/upj.0000000000000335. (PMID: 10.1097/upj.000000000000033537145731)
Huo B, Boyle A, Marfo N, Tangamornsuksan W, Steen JP, McKechnie T et al (2025) Large language models for chatbot health advice studies: a systematic review. JAMA Netw open 8(2):e2457879. https://doi.org/10.1001/jamanetworkopen.2024.57879. (PMID: 10.1001/jamanetworkopen.2024.578793990346311795331)
Menz BD, Kuderer NM, Bacchi S, Modi ND, Chin-Yee B, Hu T et al (2024) Current safeguards, risk mitigation, and transparency measures of large language models against the generation of health disinformation: repeated cross sectional analysis. BMJ (Clinical Res ed) 384:e078538. https://doi.org/10.1136/bmj-2023-078538. (PMID: 10.1136/bmj-2023-078538)
Andrikyan W, Sametinger SM, Kosfeld F, Jung-Poppe L, Fromm MF, Maas R et al (2025) Artificial intelligence-powered chatbots in search engines: a cross-sectional study on the quality and risks of drug information for patients. BMJ Qual Saf 34(2):100–109. https://doi.org/10.1136/bmjqs-2024-017476. (PMID: 10.1136/bmjqs-2024-0174763935373611874309)
Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB et al (2023) Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med 183(6):589–596. https://doi.org/10.1001/jamainternmed.2023.1838. (PMID: 10.1001/jamainternmed.2023.18383711552710148230)
Charnock D, Shepperd S, Needham G, Gann R (1999) DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Commun Health 53(2):105–111. https://doi.org/10.1136/jech.53.2.105. (PMID: 10.1136/jech.53.2.105)
Moult B, Franck LS, Brady H (2004) Ensuring quality information for patients: development and preliminary validation of a new instrument to improve the quality of written health care information. Health expectations: Int J public participation health care health policy 7(2):165–175. https://doi.org/10.1111/j.1369-7625.2004.00273.x. (PMID: 10.1111/j.1369-7625.2004.00273.x)
Silberg WM, Lundberg GD, Musacchio RA (1997) Assessing, controlling, and assuring the quality of medical information on the Internet: Caveant lector et viewor–Let the reader and viewer beware. JAMA 277(15):1244–1245. (PMID: 10.1001/jama.1997.035403900740399103351)
Rooney MK, Santiago G, Perni S, Horowitz DP, McCall AR, Einstein AJ et al (2021) Readability of patient education materials from high-impact medical journals: a 20-year analysis. J patient experience 8:2374373521998847. https://doi.org/10.1177/2374373521998847. (PMID: 10.1177/2374373521998847)
Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW et al (2023) Large language models encode clinical knowledge. Nature 620(7972):172–180. https://doi.org/10.1038/s41586-023-06291-2. (PMID: 10.1038/s41586-023-06291-23743853410396962)
Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW (2023) Large language models in medicine. Nat Med 29(8):1930–1940. https://doi.org/10.1038/s41591-023-02448-8. (PMID: 10.1038/s41591-023-02448-837460753)
Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C et al (2023) Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit health 2(2):e0000198. https://doi.org/10.1371/journal.pdig.0000198. (PMID: 10.1371/journal.pdig.0000198368126459931230)
Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG et al (2023) The future landscape of large language models in medicine. Commun Med 3(1):141. https://doi.org/10.1038/s43856-023-00370-1. (PMID: 10.1038/s43856-023-00370-13781683710564921)
Tan S, Xin X, Wu D (2024) ChatGPT in medicine: prospects and challenges: a review article. Int J Surg (London England) 110(6):3701–3706. https://doi.org/10.1097/js9.0000000000001312. (PMID: 10.1097/js9.0000000000001312)
Lee P, Bubeck S, Petro J, Benefits (2023) Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N Engl J Med 388(13):1233–1239. https://doi.org/10.1056/NEJMsr2214184. (PMID: 10.1056/NEJMsr221418436988602)
Meskó B, Topol EJ (2023) The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digit Med 6(1):120. https://doi.org/10.1038/s41746-023-00873-0. (PMID: 10.1038/s41746-023-00873-03741486010326069)
Raza MM, Venkatesh KP, Kvedar JC (2024) Generative AI and large language models in health care: pathways to implementation. NPJ Digit Med 7(1):62. https://doi.org/10.1038/s41746-023-00988-4. (PMID: 10.1038/s41746-023-00988-43845400710920625)
Alkaissi H, McFarlane SI (2023) Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 15(2):e35179. https://doi.org/10.7759/cureus.35179. (PMID: 10.7759/cureus.35179368111299939079)
Zhao YF, Bove A, Thompson D, Hill J, Xu Y, Ren Y et al (2025) Generative AI is not ready for clinical use in patient education for lower back pain patients, even with retrieval-augmented generation. AMIA joint summits on translational science proceedings AMIA joint summits on translational science. 2025:644 – 53.
Berkman ND, Sheridan SL, Donahue KE, Halpern DJ, Crotty K (2011) Low health literacy and health outcomes: an updated systematic review. Ann Intern Med 155(2):97–107. https://doi.org/10.7326/0003-4819-155-2-201107190-00005. (PMID: 10.7326/0003-4819-155-2-201107190-0000521768583)
Sørensen K, Van den Broucke S, Fullam J, Doyle G, Pelikan J, Slonska Z et al (2012) Health literacy and public health: a systematic review and integration of definitions and models. BMC Public Health 12:80. https://doi.org/10.1186/1471-2458-12-80. (PMID: 10.1186/1471-2458-12-80222766003292515)
Lenis AT, Lec PM, Chamie K, Urinary Diversion (2020) JAMA 324(21):2222. https://doi.org/10.1001/jama.2020.17604. (PMID: 10.1001/jama.2020.1760433258891)
Langford AT, Roberts T, Gupta J, Orellana KT, Loeb S (2020) Impact of the internet on patient-physician communication. Eur Urol focus 6(3):440–444. https://doi.org/10.1016/j.euf.2019.09.012. (PMID: 10.1016/j.euf.2019.09.01231582312)
Tan SS, Goonawardene N (2017) Internet health information seeking and the patient-physician relationship: a systematic review. J Med Internet Res 19(1):e9. https://doi.org/10.2196/jmir.5729. (PMID: 10.2196/jmir.5729281045795290294)
Manolitsis I, Feretzakis G, Tzelves L, Kalles D, Katsimperis S, Angelopoulos P et al (2023) Training ChatGPT models in assisting urologists in daily practice. Stud Health Technol Inform 305:576–579. https://doi.org/10.3233/shti230562. (PMID: 10.3233/shti23056237387096)
Grant Information: lzujbky-2023-ct06 Fundamental Research Funds for the Central Universities; 24ZDFA007 Gansu Province Major Science and Technology Special Project; 2023RCXM42 Gansu Province Key Talent Program; CY2024-LC-A01 Cuiying Science and Technology Innovation Program; GSWSKY2022-63 Gansu Provincial Health Industry Research Project; 2025-21-zdfy-009 Key Incubation Project Funds of the Second Hospital & Clinical Medical School, Lanzhou University
Contributed Indexing: Keywords: Artificial intelligence; Chatbot; Information quality; Patient education; Readability; Robot-assisted radical cystectomy
Entry Date(s): Date Created: 20260722 Date Completed: 20260722 Latest Revision: 20260722
Update Code: 20260722
DOI: 10.1007/s11701-026-03698-7
PMID: 42484768
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1863-2491
DOI:10.1007/s11701-026-03698-7