Can large language models like GPT-4 be trusted in clinical settings?

Can large language models like GPT-4 be trusted in clinical settings?

7. Juli 2025 um 08:48

The question of whether large language models (LLMs) such as GPT-4 can be “trusted” in clinical settings hinges not merely on their raw performance, but on a broader conception of trust that encompasses reliability, transparency, accountability, and ethical use. Below we synthesize empirical findings and normative analyses to show when and how GPT-4–class models may earn conditional trust in healthcare, and where important gaps remain.

1. Defining Trust versus Reliance Trust in clinical care implies that a tool is not only predictably reliable but also deployed in a framework that assures explainability, liability, and alignment with professional and ethical norms. Philosophers distinguish “reliance” (dependence on a system that may nonetheless be unaccountable) from “trust” (conferring authority based on perceived trustworthiness) [1]. Clinicians may rely on GPT-4 for drafting notes, but they will only trust it for decision support if its reasoning is transparent and its outputs are auditable [2].

2. Empirical Performance in Clinical Tasks 2.1 Knowledge‐based Benchmarks GPT-4 often outperforms prior models on standardized medical exams: for USMLE‐style questions, GPT-4 achieves performance comparable to or exceeding that of medical students [3], and passes the MRCS Part A surgical exam with >80% accuracy and high concordance scores [4]. However, on anesthesiology board questions GPT-4 attained ~70% correctness—below the typical 70% passing threshold for certification [5]. 2.2 Case‐based Reasoning and Ethics In a panel evaluation of medical ethics vignettes, GPT-4 scored high on clarity (4.7/5) but lower on depth (3.8/5) and acceptability (3.8/5), often missing the nuanced relational aspects of ethical dilemmas [6]. Likewise, in real‐world internal-medicine, emergency-medicine, and ethics scenarios, clinicians rated GPT-4 responses at 4.4/5 for overall utility, yet emphasized that outputs must complement, not replace, human judgment [3]. 2.3 Multimodal and EHR Interpretation Multimodal LLMs can generate radiology draft reports directly from images and requisitions, approaching supervised‐reader performance under expert oversight [7]. Prompt engineering can further boost GPT-4’s ability to extract nuanced clinical end-points—e.g., identifying arrhythmia recurrence from EHR notes rose from 64% to >91% balanced accuracy when rationale requests, structured outputs, and in-context exemplars were added [8].

3. Key Limitations and Risks 3.1 Hallucinations and Misinformation As probabilistic generators, LLMs may assert plausible but false statements (“hallucinations”) with high fluency, risking patient harm if unchecked [7][9]. 3.2 Data Bias and Equity LLMs trained on internet corpora may encode racial, gender, and socio-economic biases that perpetuate health disparities unless actively audited and mitigated [10][11]. 3.3 Knowledge Currency Without integration to live clinical databases, GPT-4’s training cutoff (early 2023) may yield outdated guidance in rapidly evolving fields. 3.4 Explainability and Accountability LLMs remain “black boxes,” challenging clinicians’ ability to trace reasoning or assign liability when errors occur [2]. Regulatory frameworks (e.g., FDA, EMA) currently lack clear pathways for certifying general-purpose LLMs as medical devices.

4. Conditions for Earning Clinical Trust 4.1 Human-in-the-Loop Integration LLM outputs should trigger clinician review, not autonomous action, with clear interfaces for acceptance, editing, or rejection [10]. 4.2 Rigorous, Domain-Specific Validation Prospective studies and randomized trials are needed to measure patient outcomes when GPT-4–enabled tools inform triage, diagnosis, or treatment planning [12]. 4.3 Continuous Monitoring and Updating Real-time performance metrics must detect drift, biases, or model degradation, with pipelines for rapid retraining or deactivation [13]. 4.4 Explainable and Auditable Architectures Embedding provenance metadata—e.g., source documents, confidence scores, chain‐of‐thought logs—can improve transparency and support post-hoc review [14]. 4.5 Ethical and Legal Oversight Co-design with ethicists, legal experts, and patient representatives ensures that deployment aligns with principles of autonomy, beneficence, non-maleficence, and justice [6][15].

5. Pragmatic Use-Cases Today The least risky domains for GPT-4 deployment are administrative and educational tasks—drafting discharge summaries, coding, patient-facing educational materials, or translation of radiology reports for non-English speakers—always under expert supervision [7][9][16]. Early adopters should focus on augmenting clinician workflows rather than supplanting diagnostic or therapeutic decision-making. Conclusion GPT-4 and similar LLMs exhibit strong capabilities in language generation, knowledge synthesis, and even multimodal reasoning, and they can relieve cognitive and administrative burdens in healthcare. Yet “trust” in the clinical sense demands more than raw accuracy—it requires transparent reasoning, up-to-date and unbiased knowledge, robust validation, and clear accountability structures. Under a human-in-the-loop paradigm, with continuous monitoring, ethical oversight, and targeted regulatory approval, LLMs can become trustworthy assistants in healthcare. Autonomous use for direct diagnosis or treatment remains premature without further prospective trials and explainability advances.

Referenzen
  1. [1]

    HATHERLEY, Joshua. Limits of trust in medical AI [preprint]. arXiv, 2020. arXiv:2503.16692. https://doi.org/10.1136/medethics-2019-105935.

  2. [2]

    JONES, Caroline; THORNTON, James; WYATT, J. Artificial intelligence and clinical decision support: Clinicians’ perspectives on trust, trustworthiness, and liability. Medical Law Review, 2023. https://doi.org/10.1093/medlaw/fwad013.

  3. [3]

    LAHAT, Adi, et al. Assessing generative pretrained transformers (GPT) in clinical decision-making: Comparative analysis of GPT-3.5 and GPT-4. Journal of Medical Internet Research, 2024. https://doi.org/10.2196/54571.

  4. [4]

    YIU, A.; LAM, K. Performance of large language models at the MRCS part a: A tool for medical education? Annals of The Royal College of Surgeons of England, 2023. https://doi.org/10.1308/rcsann.2023.0085.

  5. [5]

    KHAN, A., et al. Artificial intelligence for anesthesiology board-style examination questions: Role of large language models. Journal of cardiothoracic and vascular anesthesia, 2024. https://doi.org/10.1053/j.jvca.2024.01.032.

  6. [6]

    BALAS, M., et al. Exploring the potential utility of AI large language models for medical ethics: An expert panel evaluation of GPT-4. Journal of Medical Ethics, 2023. https://doi.org/10.1136/jme-2023-109549.

  7. [7]

    BHAYANA, Rajesh. Chatbots and large language models in radiology: A practical primer for clinical and research applications. Radiology, 2024. https://doi.org/10.1148/radiol.232756.

  8. [8]

    FENG, Ruibin, et al. Engineering of generative artificial intelligence and natural language processing models to accurately identify arrhythmia recurrence. Circulation. Arrhythmia and electrophysiology, 2024. https://doi.org/10.1161/circep.124.013023.

  9. [9]

    YANG, He S., et al. AI chatbots in clinical laboratory medicine: Foundations and trends. Clinical chemistry, 2023. https://doi.org/10.1093/clinchem/hvad106.

  10. [10]

    YU, Ping, et al. Leveraging generative AI and large language models: A comprehensive roadmap for healthcare integration. Healthcare, 2023. https://doi.org/10.3390/healthcare11202776.

  11. [11]

    LAW, Saikam; OLDFIELD, Brian; YANG, Wah. Chatgpt/gpt‐4 (large language models): Opportunities and challenges of perspective in bariatric healthcare professionals. Obesity Reviews, 2024. https://doi.org/10.1111/obr.13746.

  12. [12]

    PRESSMAN, Sophia M, et al. Clinical and surgical applications of large language models: A systematic review. Journal of Clinical Medicine, 2024. https://doi.org/10.3390/jcm13113041.

  13. [13]

    UPADHYAY, Umashankar, et al. Call for the responsible artificial intelligence in the healthcare. BMJ Health & Care Informatics, 2023. https://doi.org/10.1136/bmjhci-2023-100920.

  14. [14]

    SIVARAJKUMAR, Sonish, et al. An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: Algorithm development and validation study. JMIR Medical Informatics, 2024. https://doi.org/10.2196/55318.

  15. [15]

    FOURNIER-TOMBS, Eleonore; MCHARDY, J. A medical ethics framework for conversational artificial intelligence. Journal of Medical Internet Research, 2022. https://doi.org/10.2196/43068.

  16. [16]

    KHANNA, Praneet, et al. Artificial intelligence in multilingual interpretation and radiology assessment for clinical language evaluation (AI-MIRACLE). Journal of Personalized Medicine, 2024. https://doi.org/10.3390/jpm14090923.

7. Juli 2025 um 08:48

tlooto kann Fehler machen. Prüfe wichtige Informationen mit den Originalquellen.