Multimodal Machine Learning ApplicationsGenerative Adversarial Networks and Image SynthesisVideo Analysis and Summarization

Tingting Han, Yuxuan Gong, Sicheng Zhao, Min Tan, Zhou Yu, Hongxun Yao

2026.4.1IEEE Transactions on Affective Computing

DOI: 10.1109/taffc.2026.3652228

Abstract

Emotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we propose HEART (Hierarchical Emotion-Aligned Representation with Temporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions.

Citation format

HAN, Tingting, et al. HEART: Emotionally grounded video captioning via hierarchical emotion-aligned representation. IEEE Transactions on Affective Computing, 2026, 17(2): 1709–1720.