Open AccessMedicineComputer ScienceEducation

Tiffany H. Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, Victor Tseng

2022.12.20PLOS Digital Health

DOI: 10.1371/journal.pdig.0000198

tlooto Summary

The performance of a large language model called ChatGPT on the USMLE, which consists of three exams, suggests that large language models may have the potential to assist with medical education, and potentially, even clinical decision-making.

Abstract

We evaluated the performance of a large language model called ChatGPT on the United States Medical Licensing Exam (USMLE), which consists of three exams: Step 1, Step 2CK, and Step 3. ChatGPT performed at or near the passing threshold for all three exams without any specialized training or reinforcement. Additionally, ChatGPT demonstrated a high level of concordance and insight in its explanations. These results suggest that large language models may have the potential to assist with medical education, and potentially, even clinical decision-making.

Citation format

KUNG, Tiffany H., et al. Performance of chatgpt on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health, 2022, 2.