Multimodal Machine Learning ApplicationsAdvanced Image and Video Retrieval TechniquesVisual Attention and Saliency Detection

Xin Huang, Shilong Wang, Tong Jia, Zhimin Yuan, Runsheng Zhang, Jingjing Li

2026.2.11ACM Transactions on Multimedia Computing Communications and Applications

DOI: 10.1145/3797043

tlooto Summary

The Adaptive Co-operative Knowledge Enhancement (ACKE) method, which comprises the Uncertainty-Aware Inspire Potential (UAIP) and Adaptive Co-operative Prompt (ACP) strategies, is proposed, which constructs a prompt pool where instance-specific visual prompts are dynamically selected and projected into text prompts, which collaborate to guide modal encoders toward deep semantic consensus.

Abstract

With the rapid growth of internet multimedia data, cross-modal retrieval techniques have garnered significant attention. Given the inherent complexity and non-intuitive nature of cross-modal relationships, tuning pre-trained Large Multimodal Models (LMMs) with cross-modal data has become a mainstream approach. However, cross-modal data commonly exhibit inter-modal information asymmetry and intra-modal distribution diversity. Faced with these challenges, existing paradigms tend to learn ambiguous and asymmetric cross-modal associations, which introduce semantic noise. In addition, their limited adaptability to the high diversity of real-world content further hinders optimal retrieval performance. To address these challenges, this paper proposes the Adaptive Co-operative Knowledge Enhancement (ACKE) method, which comprises the Uncertainty-Aware Inspire Potential (UAIP) and Adaptive Co-operative Prompt (ACP) strategies. UAIP utilizes generative LMMs to generate multi-perspective descriptions that enrich semantic information, while employing Dempster-Shafer Theory (DST) to quantify their semantic uncertainty and adjust contribution weights, reducing inaccurate relational mappings and balancing information asymmetry. ACP constructs a prompt pool where instance-specific visual prompts are dynamically selected and projected into text prompts, which collaborate to guide modal encoders toward deep semantic consensus, thus mitigating alignment bias from intra-modal distribution diversity and improving accuracy. Extensive experiments are conducted on two widely used datasets, Flickr30K and MS-COCO, demonstrating the effectiveness of our proposed method. The code is available at https://github.com/nynu-BDAI/ACKE.

Citation format

HUANG, Xin, et al. Adaptive co-operative prompting and uncertainty-aware implicit knowledge enhancement for cross-modal retrieval. ACM Transactions on Multimedia Computing Communications and Applications, 2026, 22(5): 1–26.