MedicineComputer Science

Amirreza Izadi, H. Mosavari, Ali Hosseininasab, Ali Jaliliyan, Arzhang Jafari, Mohammadhosein Akhlaghpasand, Aghil Rostami, M. Moradi-Lakeh, Foolad Eghbali

2026.1.1Journal of Obesity

DOI: 10.1155/jobe/2376530

tlooto Summary

AI chatbots generated accurate and comprehensive answers to common bariatric patient questions, suggesting promise as a scalable aid for patient education, however, readability often exceeds recommended levels, performance varies by model, occasional inaccuracies occur, and medicolegal considerations remain unresolved.

Abstract

Background The global obesity epidemic challenges health systems, driving people to seek metabolic and bariatric surgery (MBS), especially laparoscopic sleeve gastrectomy (LSG). Many MBS centers have limited resources for patient education, creating knowledge gaps that lead patients to search online. AI chatbots, such as ChatGPT, can provide reliable medical information, though concerns about accuracy and completeness remain. Methods The study involved four fellowship‐trained minimally invasive surgeons (MISs), nine fellows (MIFs), and two general practitioners (GPs) in the MBS multidisciplinary team from March 1, 2024, to March 30, 2024. Seven AI chatbots were selected, including ChatGPT 3.5 and 4, Bard, Bing, Claude, Llama, and Perplexity, based on their public availability on December 1, 2023. Forty patient questions regarding LSG were sourced from social media, MBS organizations, and online forums. Experts and chatbots answered these questions, with their responses evaluated for accuracy and comprehensiveness on a 5‐point scale. Statistical analyses compared groups’ performance. Results Chatbots demonstrated a higher overall performance score (2.55 ± 0.95) compared to the expert group (1.92 ± 1.32, p < 0.001). Among chatbots, ChatGPT‐4 achieved the highest performance (2.94 ± 0.24), while Llama had the lowest (2.15 ± 1.23). Expert group scores were highest for MISs (2.36 ± 1.09), followed by GPs (1.90 ± 1.36) and MIFs (1.75 ± 1.36). The readability of chatbot responses was assessed using Flesch–Kincaid scores, revealing that most responses required reading levels between the 11th grade and college level. Furthermore, chatbots exhibited fair reliability and reproducibility in response consistency, with ChatGPT‐4 showing the highest test–retest reliability. Conclusion AI chatbots generated accurate and comprehensive answers to common bariatric patient questions, suggesting promise as a scalable aid for patient education. However, readability often exceeds recommended levels, performance varies by model, occasional inaccuracies occur, and medicolegal considerations remain unresolved. Accordingly, chatbots should complement clinician counseling, and future work should improve readability and reliability and evaluate real‐world safety and impact.

Citation format

IZADI, Amirreza, et al. Patient education in bariatric surgery: Can artificial intelligence–based chatbots bridge the knowledge gap? Journal of Obesity, 2026, 2026(1): 2376530.