Timothy Meinert, Anna Koufakou
2026.5.6Proceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS
Abstract
Generative large language models (LLMs) are often assumed to outperform earlier transformer-based encoders across NLP tasks, yet this has not been adequately tested for emotion classification. Using a recently introduced multi-dataset emotion benchmark, we compare a Llama-based generative model with previously reported results from a fine-tuned RoBERTa classifier. The zero-shot LLM consistently underperforms while few-shot prompting substantially improves LLM performance for several datasets. These findings challenge the assumption that LLMs universally surpass older transformers and highlight the continued relevance of fine-tuned models for emotion classification. At the same time, they show that few-shot prompting can unlock competitive LLM performance without the need for task-specific training but not for all datasets.
Citation format
MEINERT, Timothy; KOUFAKOU, Anna. Do LLMs outperform fine-tuned transformers in emotion classification? Proceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS, 2026, 39(1).