Claudia Buder, Nina-Sophie Fritsch, Chiara Osorio-Krauter, Aaron Philipp, Roland Verwiebe, Sarah Weißmann
Abstract
This paper examines context-sensitive annotation tasks performed by OpenAI’s o3 reasoning model using content creator profiles on YouTube. It explores the inherently ambiguous and socially and culturally embedded task of annotating race. Analysing 500 annotations generated by ChatGPT, we first examine performance metrics and benchmark results against a human-annotated dataset. We further conduct a thematic analysis of the justifications provided by the model, as users increasingly rely on them making them important for understanding how the model frames social practices. Analyzing justifications also exposes inconsistencies, biases, and classification errors that remain invisible in aggregate performance metrics alone. Our findings show that despite efficiency and scalability, ChatGPT's limited cultural understanding and lack of critical reflexivity constrain its performance in complex annotation tasks. We argue that the annotation of sensitive social characteristics requires reflexive scientific practices and potentially hybrid annotation strategies to mitigate bias and preserve contextual integrity in academic research.
Citation format
BUDER, Claudia, et al. Efficiency vs. understanding: A critical examination of chatgpt’s performance in context-sensitive annotation tasks. International Journal of Sociology, 2026.