MedicineComputer Science

Ou Huan, Zhen Wang

2026.3.3ARCHIVES OF GYNECOLOGY AND OBSTETRICS

DOI: 10.1007/s00404-026-08358-7

tlooto Summary

Under constrained non-reasoning prompts, even next-generation AI chatbots demonstrate unsatisfactory performance in gynecology, highlighting the critical need for "chain-of-thought" prompting and human expert oversight.

Abstract

Purpose As artificial intelligence (AI) models evolve into their next generations, their application in specialized medical fields requires rigorous validation. While large language models (LLMs) have shown promise in general medicine, their reliability in complex gynecological clinical reasoning remains under-explored. This pilot study aimed to comparatively assess the knowledge retention, safety, and reasoning limitations of advanced AI chatbots in gynecology using a constrained zero-shot multiple-choice question (MCQ) format.

Citation format

HUAN, Ou; WANG, Zhen. Performance of next-generation AI chatbots in gynecological knowledge assessment: A comparative pilot study of chatgpt-5, gemini-3, deepseek-v3.2, and claude-4.5-opus. ARCHIVES OF GYNECOLOGY AND OBSTETRICS, 2026, 313.