Ruifan Li, Mingcong Lu, Pengyue Lin, Zhihan Yu, Zhanyu Ma
2026.1.1IEEE MULTIMEDIA
Abstract
Scene knowledge referring expression comprehension is an emerging multimedia reasoning task. It requires the model to perform joint reasoning over referring expressions and scene knowledge to locate target objects. However, the complexity of scene knowledge and redundant information could cause interference with models. To this end, we propose a data simplification scheme. We leverage the understanding capability of large language models to simplify the complex scene knowledge. Then, we can filter out irrelevant descriptions and keep those relevant to the target objects involved by the referring expression. Furthermore, we propose a scene knowledge reasoning network (SKRN). Our SKRN extracts features from both referring expressions and scene knowledge and employs an attention mechanism to fully utilize them for reasoning. This enhances the model’s ability to handle scene knowledge and ultimately improves localization accuracy. Experimental results on the benchmark dataset demonstrate the effectiveness of our data simplification scheme and the proposed SKRN.
Citation format
LI, Ruifan, et al. Improving scene knowledge referring expression comprehension with large language models. IEEE MULTIMEDIA, 2026, 33(1): 72–80.