Yang Yang, Jiankun Yang, Haibo Lu, Huaping Liu, Wen Gao
2026.6.1IEEE Systems Journal
Abstract
In this article, we propose a multimodal robotic agent that addresses the limitations of passive perception and fixed-model deployment through a flexible large-model architecture. This architecture enables dynamic selection of multimodal large language models based on computational constraints. Building on this foundation, we introduce a logic-guided active perception strategy that decides which skills (e.g., knock and weigh) to employ based on intermediate reasoning, rather than exhaustively executing all possible actions. Our work focuses on cohesive skill integration within a unified control loop, optimizing both perception and action. This allows the agent to strategically probe objects’ visual, auditory, tactile, and weight attributes for accurate material inference and robust task completion. Extensive evaluations in the simulation environment highlight the efficiency and adaptability of our method, especially in active perception and latent information inference for robotic systems.
Citation format
YANG, Yang, et al. Active perception strategies for multimodal integration and latent information reasoning in robotic systems. IEEE Systems Journal, 2026, 20(2): 427–435.