Topic ModelingSentiment Analysis and Opinion MiningAI in Service Interactions
DOI: 10.1017/s0890060426100250

Abstract

Abstract This study investigates the use of large language models (LLMs) to classify question utterances within verbal design protocols according to Eris’ (2004) taxonomy. We evaluate the performance of two proprietary LLMs – OpenAI’s GPT-4.1 and Anthropic’s Claude Sonnet 4.5 – across experiments designed to assess classification accuracy, sensitivity to prompt configuration and in-context learning (ICL), and generalization across datasets and models. Using two human-coded datasets of differing size and quality, we measure alignment between LLM-generated labels and human judgments at both question category and subcategory levels. Results show that both LLMs achieved moderate to strong alignment rates at the category level (up to 85.7% for GPT-4.1 and 82.9% for Claude Sonnet 4.5), with substantially lower alignment at the more granular subcategory level. Performance differences across prompt configurations and ICL conditions were small, indicating robust generalization across datasets and transferability of prompt designs. While these results suggest that LLMs can effectively support scalable question classification, human judgment and oversight remain essential. Future research should explore the development and evaluation of alternative hybrid human–LLM workflows in protocol analysis, as well as the use of smaller or open-source models to address data privacy concerns.

Citation format

SAKIB, Ahmed Shahriar; HURST, Ada; SAFAYENI, F. Evaluating the performance of large language models in taxonomic classification of questions in verbal protocols of design. AI EDAM-ARTIFICIAL INTELLIGENCE FOR ENGINEERING DESIGN ANALYSIS AND MANUFACTURING, 2026, 40.