Emotion and Mood RecognitionSentiment Analysis and Opinion MiningSocial Robot Interaction and HRI

Zhongshi Xu

2026.1.1Alexandria Engineering Journal

DOI: 10.1016/j.aej.2025.11.043

Abstract

Emotion generation plays a key role in multimodal affective computing, such as in intelligent customer service and virtual assistants. However, most existing methods rely on a single modality or simple modal fusion, failing to fully capture the complex relationships between multimodal information. This results in emotional responses that lack consistency and diversity. To address these issues, we propose an interactive emotion-guided generation method (MMEG) based on multimodal data fusion. MMEG combines graph convolutional networks (GCN) and cross-modal attention mechanisms to capture complex modality dependencies. It also employs generative adversarial networks (GANs) to enhance the quality of generated responses. Experimental results on the IEMOCAP and MELD datasets show that MMEG outperforms existing methods in emotion recognition accuracy, generation quality, and response diversity. The model achieves superior performance in key metrics such as ROUGE-L, BLEU, and F1-score, while also improving emotional consistency. This method offers an effective solution for multimodal emotion generation with broad applications in affective computing and intelligent interaction.

Citation format

XU, Zhongshi. Interactive emotion-guided generation with cross-modal attention and graph convolutional networks. Alexandria Engineering Journal, 2026, 134: 197–209.