Computer ScienceEducation

Congcong Xie, Di Wang, Quan Wang, Xiao Liang, Ruyi Liu, Qiguang Miao

2026.1.1IEEE Transactions on Learning Technologies

DOI: 10.1109/tlt.2026.3656606

Abstract

Real-time feedback in online education relies on the automated assessment of student engagement. Current research primarily evaluates engagement through the analysis of students’ behavioral performance in classroom videos. Although existing methods consider temporal and multimodal cues, they predominantly rely on visual information and insufficiently model the cognitive rhythm of attention. Moreover, several approaches are limited to engagement classification and fail to offer continuous quantitative assessment. To address these limitations, we propose M-LATTE method, short for multimodal latent attention trends and time-cyclic engagement modeling. The model extracts features from visual, audio, and textual modalities and employs a cross-modal attention mechanism to achieve effective multimodal fusion, thus avoiding solely relying on visual information. The fused temporal data are decomposed into long-term trends and time-cyclic fluctuations to capture cognitive rhythm characteristics. A smoothness constraint and a variational lower bound are introduced to suppress transient disturbances, ensuring stable evaluation. Experimental results show significant performance improvements over the baseline: on the RoomReader dataset, mean squared error decreases from 0.1912 to 0.0969. Furthermore, our method achieves a competitive classification accuracy of 61.37% on the dataset for affective states in e-environments (DAiSEE) dataset.

Citation format

XIE, Congcong, et al. Multimodal latent temporal modeling for continuous engagement assessment in online education. IEEE Transactions on Learning Technologies, 2026, 19: 21–34.