W. Almutairi
2026.4.25Saudi Endodontic Journal
Abstract
Artificial intelligence (AI) chatbots are increasingly used by the public to obtain medical and dental information, including inquiries related to endodontic surgery. Evaluating the accuracy of these platforms is essential to ensure reliable patient education. This study aimed to assess and compare the reliability of three advanced AI chatbot platforms – ChatGPT-5, DeepSeek, and Grok in responding to patient-style questions related to surgical endodontic treatment. A set of 16 patient-centered questions covering four domains (indications, procedure, postoperative care, and complications) was presented to each AI platform. Responses were independently evaluated by two board-certified specialists using a 5-point Likert-type scale (1 = strongly disagree to 5 = strongly agree) reflecting alignment with evidence-based surgical endodontic standards. Descriptive statistics were used to summarize the distribution of ratings, and a Chi-square test was prespecified for potential comparative analysis. All platforms demonstrated consistently high reliability. ChatGPT-5 and DeepSeek each received 15 ratings of 5 and 1 rating of 4, while Grok received 14 ratings of 5 and 2 ratings of 4. No responses from any platform were rated as neutral or negative. Because all responses were rated “agree” or “strongly agree,” no inferential statistical testing was performed. All tested AI chatbot platforms performed comparably in delivering clinically relevant and accurate responses to patient-style endodontic questions. These findings suggest that modern AI chatbots may serve as valuable adjuncts in dental communication and education, particularly in enhancing patient understanding and engagement.
Citation format
ALMUTAIRI, W. Reliability of three artificial intelligence chatbots in responding to patient questions about surgical endodontics: An observational study. Saudi Endodontic Journal, 2026, 16(2): 266–270.