MedicineComputer Science

Bita Momenaei, Taku Wakabayashi, A. Shahlaee, Asad F. Durrani, Saagar A. Pandit, Kristine Wang, Hana A. Mansour, Robert M. Abishek, David Xu, J. Sridhar, Y. Yonekawa, Ajay E. Kuriyan

2023.6.1Ophthalmology Retina

DOI: 10.1016/j.oret.2023.05.022

tlooto Summary

Evaluating the appropriateness and readability of the medical knowledge provided by ChatGPT-4, an artificial-intelligence-powered conversational search engine regarding common vitreoretinal surgeries for retinal detachments, macular holes, and epiretinal membranes found most of the answers provided were consistently appropriate.

Abstract

OBJECTIVE To evaluate the appropriateness and readability of the medical knowledge provided by ChatGPT-4, an artificial-intelligence-powered conversational search engine regarding common vitreoretinal surgeries for retinal detachments, macular holes, and epiretinal membranes.

DESIGN Retrospective cross-sectional study.

SUBJECTS This study does not involve any human participants.

METHODS We created lists of common questions about the definition, prevalence, visual impact, diagnostic methods, surgical and non-surgical treatment options, postoperative information, surgery-related complications, and visual prognosis of retinal detachment, macular hole, and epiretinal membrane, and asked each question three times on the online ChatGPT-4 platform. The data for this cross-sectional study were recorded on April 25, 2023. Two independent retina specialists graded the appropriateness of the responses. Readability was assessed using Readable, an online readability tool.

MAIN OUTCOME MEASURES The "appropriateness" and "readability" of the answers generated by ChatGPT-4 bot.

RESULTS Responses were consistently appropriate in 84.6% (33/39), 92% (23/25), and 91.7% (22/24) of the questions related to retinal detachment, macular hole, and epiretinal membrane, respectively. Answers were inappropriate at least once in 5.1% (2/39), 8% (2/25), and 8.3% (2/24) of the respective questions. The average Flesch Kincaid Grade Level and Flesch Reading Ease Score were 14.1±2.6 and 32.3±10.8 for retinal detachment, 14±1.3 and 34.4±7.7 for macular hole, and 14.8±1.3 and 28.1±7.5 for epiretinal membrane. These scores indicate that the answers are difficult or very difficult to read for the average lay person and college graduation would be required to understand the material.

CONCLUSIONS Most of the answers provided by ChatGPT-4 were consistently appropriate. However, ChatGPT and other natural language models in their current form are not a source of factual information. Improving the credibility and readability of responses, especially in specialized fields such as medicine, is a critical focus of research. Patients, physicians, and laypersons should be advised of the limitations of these tools for eye- and health-related counseling.

Citation format

MOMENAEI, Bita, et al. Appropriateness and readability of chatgpt-4 generated responses for surgical treatment of retinal diseases. Ophthalmology Retina, 2023.