Ou Huan, Zhen Wang
tlooto Summary
Under constrained non-reasoning prompts, even next-generation AI chatbots demonstrate unsatisfactory performance in gynecology, highlighting the critical need for "chain-of-thought" prompting and human expert oversight.
Abstract
Purpose As artificial intelligence (AI) models evolve into their next generations, their application in specialized medical fields requires rigorous validation. While large language models (LLMs) have shown promise in general medicine, their reliability in complex gynecological clinical reasoning remains under-explored. This pilot study aimed to comparatively assess the knowledge retention, safety, and reasoning limitations of advanced AI chatbots in gynecology using a constrained zero-shot multiple-choice question (MCQ) format.
Citation format
HUAN, Ou; WANG, Zhen. Performance of next-generation AI chatbots in gynecological knowledge assessment: A comparative pilot study of chatgpt-5, gemini-3, deepseek-v3.2, and claude-4.5-opus. ARCHIVES OF GYNECOLOGY AND OBSTETRICS, 2026, 313.