Open AccessMedicineComputer Science

F. Antaki, Samir Touma, D. Milad, J. El-Khoury, R. Duval

2023.1.24Ophthalmology Science

DOI: 10.1016/j.xops.2023.100324

tlooto Summary

ChatGPT's performance varied across subspecialties, with the best results in general medicine and the worst in neuro-ophthalmology and ophthalmic pathology and intraocular tumors, suggesting that specialising LLMs through domain-specific pre-training may be necessary to improve their performance in ophthalmology sub specialties.

Abstract

We tested the accuracy of ChatGPT, a large language model (LLM), in the ophthalmology question-answering space using two popular multiple choice question banks used for the high-stakes Ophthalmic Knowledge Assessment Program (OKAP) exam. The testing sets were of easy-to-moderate difficulty and were diversified, including recall, interpretation, practical and clinical decision-making problems. ChatGPT achieved 55.8% and 42.7% accuracy in the two 260-question simulated exams. Its performance varied across subspecialties, with the best results in general medicine and the worst in neuro-ophthalmology and ophthalmic pathology and intraocular tumors. These results are encouraging but suggest that specialising LLMs through domain-specific pre-training may be necessary to improve their performance in ophthalmic subspecialties.

Citation format

ANTAKI, F., et al. Evaluating the performance of chatgpt in ophthalmology. Ophthalmology Science, 2023, 3.