Intelligent Tutoring Systems and Adaptive LearningWriting and Handwriting EducationTopic Modeling

Jie Zhang, Chunhua Ma

2026.2.1Chinese Journal of Applied Linguistics

DOI: 10.1515/cjal-2026-0105

tlooto Summary

Results show that AI exhibits high self-consistency and adapts effectively to different scoring roles (e.g., teacher vs. highstakes rater), but highlight the need for further calibration to align with educational standards.

Abstract

Abstract This study examines the reliability and validity of AI-generated scoring for continuation writing tasks. By comparing GPT-4 with eight experienced human raters across 21 student responses, it evaluates AI’s consistency, severity, and alignment with human scoring criteria. Results show that AI exhibits high self-consistency and adapts effectively to different scoring roles (e.g., teacher vs. highstakes rater). However, AI scores were more lenient than human raters and demonstrated divergent evaluation focuses—prioritizing narrative coherence and emotional depth, while teachers emphasized linguistic accuracy and richness of detail. The findings suggest AI’s potential as a supplementary assessment tool, offering rapid, holistic feedback, but highlight the need for further calibration to align with educational standards. Implications include exploring hybrid evaluation models that leverage the strengths of both AI and human raters to achieve more equitable, efficient, and pedagogically meaningful writing assessments.

Citation format

ZHANG, Jie; MA, Chunhua. Exploring the reliability and validity of AI scoring in the continuation writing task. Chinese Journal of Applied Linguistics, 2026, 49(1): 133–150.