Multimodal Machine Learning ApplicationsEmbodied and Extended CognitionSpeech and dialogue systems

Yang Yang, Jiankun Yang, Haibo Lu, Huaping Liu, Wen Gao

2026.6.1IEEE Systems Journal

DOI: 10.1109/jsyst.2026.3680039

Abstract

In this article, we propose a multimodal robotic agent that addresses the limitations of passive perception and fixed-model deployment through a flexible large-model architecture. This architecture enables dynamic selection of multimodal large language models based on computational constraints. Building on this foundation, we introduce a logic-guided active perception strategy that decides which skills (e.g., knock and weigh) to employ based on intermediate reasoning, rather than exhaustively executing all possible actions. Our work focuses on cohesive skill integration within a unified control loop, optimizing both perception and action. This allows the agent to strategically probe objects’ visual, auditory, tactile, and weight attributes for accurate material inference and robust task completion. Extensive evaluations in the simulation environment highlight the efficiency and adaptability of our method, especially in active perception and latent information inference for robotic systems.

Citation format

YANG, Yang, et al. Active perception strategies for multimodal integration and latent information reasoning in robotic systems. IEEE Systems Journal, 2026, 20(2): 427–435.