GPT-4와 같은 언어모델은 실제 임상에서 얼마나 신뢰할 수 있을까?
GPT-4와 같은 언어모델은 실제 임상에서 얼마나 신뢰할 수 있을까?
GPT-4와 같은 언어모델은 실제 임상에서 얼마나 신뢰할 수 있을까?
GPT-4와 같은 언어모델은 실제 임상에서 얼마나 신뢰할 수 있을까?
GPT-4와 같은 대규모 언어모델(LLM)은 자연어 이해와 생성 능력 측면에서 획기적인 발전을 이뤘으나, 실제 임상 현장에 적용하기 위해서는 정확성·안정성·윤리적·법적 책임 등의 측면에서 면밀한 검증이 필요하다. 아래에서는 주요 성능 지표와 한계, 활용 방안 및 향후 과제를 정리하였다.
임상적 성능 및 정확도 • GPT-4는 외과 전문의 국가고사 문제(한국 일반외과)를 푸는 데 76.4%의 정답률을 보여 GPT-3.5(46.8%) 대비 유의미하게 높은 성능을 기록했다(P < 0.001)[1]. • 방사선과 영역에서는 텍스트·영상 멀티모달 LLM이 판독·보고서 생성에 있어 전공의 수준에 근접하는 결과를 내며, 라디올로지 워크플로우 전 단계(요청부터 최종 리포트 작성)에 보조 역할을 할 수 있다[2]. • 연구용 의료 기록에서 핵심 정보를 추출해 생존 예측 모델에 통합한 사례에서, GPT-4는 수기 추출 대비 94–97%의 데이터 추출 정확도를 보였고(AUC ≈ 0.87–0.88), 모델 성능 저하가 크지 않았다[3]. • 임상화학 퀴즈(109문항)에서 GPT-4o는 81.7% 정답률을 달성, 다른 LLM 대비 우수함을 보이며 전문 분야 지식 응용 가능성을 시사했다[4].
허위 정보(hallucination) 및 안전성 문제 • LLM은 실제로 존재하지 않는 근거를 “그럴듯하게” 생성할 수 있어, 진단·치료 제안 시 치명적인 오류를 유발할 위험이 있다[2][5]. • 의료 데이터 특유의 복잡성(불완전·불일치·누락 정보)을 모델이 온전히 이해하지 못하면 오진 가능성이 높아진다.
데이터·학습 한계 • 사전학습(pre-training)은 공개 텍스트·논문·웹문서를 기반으로 하며, 개별 환자의 전자의무기록(EMR)이나 최신 가이드라인·실시간 검사결과는 반영되지 않을 수 있다. • 프라이버시 보호를 위해 RAG(retrieval-augmented generation)나 온디바이스 미세조정(fine-tuning)이 필요하지만, 임상 적용을 위한 규제·표준화가 아직 미비하다[6].
윤리·법적 책임 • LLM이 독자적으로 진단·치료 결정을 내릴 경우, 의료 사고가 발생했을 때 법적 책임 소지가 불분명하다. • 환자 동의 획득, 데이터 익명화, 알고리즘 투명성 확보 등 윤리적 규율(의료윤리·개인정보보호법 등)을 준수해야 한다[7].
보조적 도구로서의 활용 • 임상 의사결정 지원(CDS): 증례 리뷰, 감별진단 목록 확장, 가이드라인 서지 검색 등에 활용 가능하다[8][9]. • 문서 작업 자동화: EMR 요약, 보고서 초안 작성, 환자 안내문 생성, 의료윤리 자문 지원 등으로 의료진 부담을 경감시킨다[7][10]. • 임상시험 대상자 선별: 심부전 임상시험 프리스크리닝에서 AI-모델 사용 시 수작업 대비 효율을 높였다는 임상시험 결과가 있다[11].
신뢰성 향상을 위한 전략 • 도메인 특화 미세조정: 방사선·병리·약물학 전용 데이터셋으로 추가 학습하여 전문성 강화 필요[12][13]. • 인간-중심 협업(Co-design): 임상의·윤리전문가·개발자가 공동으로 시스템을 설계·검증해야 한다[6]. • 지속적 모니터링·검증: 실제 사용 환경에서 정량적 성능 평가와 정성적 사용자 피드백을 병행해야 한다.
결론적으로, GPT-4 같은 LLM은 임상 지원 도구로서 잠재력이 크나 단독 임상 의사결정 수단으로는 아직 부족하다. 정확도·안정성·윤리적 합의·법적 프레임워크가 충분히 확보된 뒤에야, 의료진과 협업하는 형태로 실질적인 임상 효용을 달성할 수 있을 것이다.
OH, N.; CHOI, G.; LEE, W. Chatgpt goes to the operating room: Evaluating GPT-4 performance and its potential in surgical education and training in the era of large language models. Annals of Surgical Treatment and Research, 2023. https://doi.org/10.4174/astr.2023.104.5.269.
BHAYANA, Rajesh. Chatbots and large language models in radiology: A practical primer for clinical and research applications. Radiology, 2024. https://doi.org/10.1148/radiol.232756.
SUN, Di, et al. Outcome prediction using multi-modal information: Integrating large language model-extracted clinical information and image analysis. Cancers, 2024. https://doi.org/10.3390/cancers16132402.
HEO, Won Young; PARK, H. Assessment of large language models in medical quizzes for clinical chemistry and laboratory management: Implications and applications for healthcare artificial intelligence. Scandinavian Journal of Clinical and Laboratory Investigation, 2025. https://doi.org/10.1080/00365513.2025.2466054.
OMIYE, Jesutofunmi A., et al. Large language models in medicine: The potentials and pitfalls [preprint]. arXiv, 2023. arXiv:2309.00087. https://doi.org/10.7326/m23-2772.
YU, Ping, et al. Leveraging generative AI and large language models: A comprehensive roadmap for healthcare integration. Healthcare, 2023. https://doi.org/10.3390/healthcare11202776.
BALAS, M., et al. Exploring the potential utility of AI large language models for medical ethics: An expert panel evaluation of GPT-4. Journal of Medical Ethics, 2023. https://doi.org/10.1136/jme-2023-109549.
EGUIA, Hans, et al. Clinical decision support and natural language processing in medicine: Systematic literature review. Journal of Medical Internet Research, 2024. https://doi.org/10.2196/55315.
ZHANG, Kuo, et al. Revolutionizing health care: The transformative impact of large language models in medicine. Journal of Medical Internet Research, 2024. https://doi.org/10.2196/59069.
ABDULNAZAR, Akhila, et al. Large language models for clinical text cleansing enhance medical concept normalization. IEEE Access, 2024. https://doi.org/10.1109/access.2024.3472500.
UNLU, Ozan, et al. Manual vs AI-Assisted prescreening for trial eligibility using large language models-a randomized clinical trial. JAMA, 2025. https://doi.org/10.1001/jama.2024.28047.
KIM, Kiduk, et al. Updated primer on generative artificial intelligence and large language models in medical imaging for medical professionals. Korean Journal of Radiology, 2024. https://doi.org/10.3348/kjr.2023.0818.
HSU, Joy C, et al. Applications of advanced natural language processing for clinical pharmacology. Clinical Pharmacology & Therapeutics, 2023. https://doi.org/10.1002/cpt.3161.
tlooto can make mistakes. Check important information against the original sources.