Jie Zhang, Chunhua Ma
tlooto Summary
Results show that AI exhibits high self-consistency and adapts effectively to different scoring roles (e.g., teacher vs. highstakes rater), but highlight the need for further calibration to align with educational standards.
Abstract
Abstract This study examines the reliability and validity of AI-generated scoring for continuation writing tasks. By comparing GPT-4 with eight experienced human raters across 21 student responses, it evaluates AI’s consistency, severity, and alignment with human scoring criteria. Results show that AI exhibits high self-consistency and adapts effectively to different scoring roles (e.g., teacher vs. highstakes rater). However, AI scores were more lenient than human raters and demonstrated divergent evaluation focuses—prioritizing narrative coherence and emotional depth, while teachers emphasized linguistic accuracy and richness of detail. The findings suggest AI’s potential as a supplementary assessment tool, offering rapid, holistic feedback, but highlight the need for further calibration to align with educational standards. Implications include exploring hybrid evaluation models that leverage the strengths of both AI and human raters to achieve more equitable, efficient, and pedagogically meaningful writing assessments.
Citation format
ZHANG, Jie; MA, Chunhua. Exploring the reliability and validity of AI scoring in the continuation writing task. Chinese Journal of Applied Linguistics, 2026, 49(1): 133–150.