Yiren Li, Mingyu Lv, Liang Cheng, Yongxu Xie, Qiang Gu, Haozheng He, Zhuguan Chen, Duopei Fang, Xiang Zhou
2026.2.16EUROPEAN JOURNAL OF RADIOLOGY
tlooto Summary
Among the four evaluated LLM agents, Gemini 2.5 Flash and DeepSeek demonstrated superior clinical accuracy, comprehensiveness, and usability in spine-related decision support, and support the potential of domain-adapted RAG agents to enhance evidence-based spinal care by providing accurate and comprehensive decision support.
Abstract
BACKGROUND Clinical decision-making in spinal and spinal cord diseases requires a comprehensive assessment of imaging findings, neurological status, bone integrity, and patient-centered goals. Recently, the emergence of Large Language Models (LLMs) has provided new tools for intelligent decision support; however, their reliability and clinical interpretability remain to be systematically evaluated.
METHODS We propose a novel retrieval augmented generation (RAG)-based framework specifically tailored for spinal disease decision support and systematically evaluated four advanced LLM agents (Gemini 2.5 Flash, DeepSeek, GPT-4o, GPT-4o-mini). The framework integrates domain-specific prompting, structured response formatting, and evidence citation tracking.We used 200 real-world spinal cases, each involving diagnostic, therapeutic, and follow-up tasks. Five spine surgeons independently evaluated model outputs using an 11-dimension rubric; each dimension rated on a 5-point Likert scale to evaluate both clinical and technical performance. Inter-group differences were analyzed using the Kruskal-Wallis and Dunn's tests (P < 0.05), with radar plots used for multidimensional visualization.
RESULTS The proposed expert-evaluated framework enables a comprehensive, real-case-based comparison of four LLM agents. Gemini 2.5 Flash achieved the highest overall score (49.25 ± 2.88), significantly outperforming DeepSeek (47.49 ± 3.34, P < 0.001), GPT-4o (45.59 ± 3.89, P < 0.001), and GPT-4o-mini (41.23 ± 5.96, P < 0.001). It demonstrated leading performance particularly in humanistic care, follow-up suggestion, and test recommendation. DeepSeek showed superior capability in differential completeness (mean = 4.76), significantly outperforming the other three models (P < 0.001). Although GPT-4o-mini demonstrated stable system performance (mean = 4.61), it underperformed in core clinical reasoning dimensions. These findings reveal substantial inter-model variability in spine-specific clinical reasoning, an aspect often overlooked in prior non-benchmark LLM evaluations.
CONCLUSION Among the four evaluated LLM agents, Gemini 2.5 Flash and DeepSeek demonstrated superior clinical accuracy, comprehensiveness, and usability in spine-related decision support. These findings support the potential of domain-adapted RAG agents to enhance evidence-based spinal care by providing accurate and comprehensive decision support. Future research should focus on multimodal integration data (e.g., imaging and clinical notes) and conducting prospective validation in real-world clinical environments.
Citation format
LI, Yiren, et al. Explainable and evidence-linked recommendations for spine surgery via a retrieval-augmented LLM agent. EUROPEAN JOURNAL OF RADIOLOGY, 2026, 197: 112734.