고상훈 (Sang Hun Ko), 안현철 (Hyunchul Ahn)
2024지식경영연구
tlooto Summary
LLM-based fake news detection shows limitations but potential in Korean environment, with 59.8% accuracy and improvement to 62.1% with advanced prompt design.
초록
Fake news spreads rapidly through digital platforms and social media and has become an important problem that negatively impacts social trust and public discourse. Recent advances in Large Language Models (LLMs) have opened up new possibilities for natural language processing techniques, and their use is also contributing to solving important social problems such as fake news detection. This study aims to analyze the possibilities and limitations of LLM-based fake news detection in the Korean environment, focusing on misinformation among the different types of fake news. To this end, we constructed a benchmark dataset based on 500 Korean news articles collected from SNU FactCheck. Further, we designed extracted and generated datasets by applying the article summarization method. This study centered on three research questions: (1) is LLM-based fake news detection effective in the Korean environment? (2) which summarization method is effective in detecting fake news using LLM, and (3) what is the optimal way to improve detection performance? The results showed that LLM-based detection accuracy in the Korean dataset was lower (59.8%) compared to English-focused studies. In addition, detection performance in experiments using summarized text decreased as sentences became shorter, and generated summaries performed slightly better than extracted summaries. Finally, we improved the detection accuracy slightly (62.1%) by introducing an improved prompt that reflects the seven reasons for fake news detection. This study extends the LLM-based fake news detection research in the Korean environment. It demonstrates that an advanced prompt design and an approach considering contextual factors can improve detection performance. The findings suggest the feasibility of utilizing LLM not only for fake news detection but also in various linguistic and cultural contexts, which have important implications for future research and practical applications.
인용 형식
고상훈; 안현철. 대규모 언어 모델을 활용한 한국어 가짜뉴스 탐지: 한계와 가능성. 지식경영연구, 2024, 25(4): 113–127.