Multimodal Machine Learning ApplicationsDomain Adaptation and Few-Shot LearningAdvanced Image and Video Retrieval Techniques

Jiehao Xue, Xin Xu, Zilong Zhao, Hengda Zhu, Lin Cong, Fangling Pu

2026.12.19IEEE Geoscience and Remote Sensing Letters

DOI: 10.1109/lgrs.2025.3646482

Abstract

Visual grounding in remote sensing (RSVG) images aims to localize objects referenced by natural language descriptions. While most prior work focuses on single-target grounding via cross-modal feature alignment, many real-world scenarios require the simultaneous localization of multiple objects described with more complex language. We emphasize the importance of this multitarget setting, which captures richer linguistic dependencies and spatial relations. Progress has been limited by the difficulty of repurposing existing datasets, as multitarget visual grounding (VG) demands far more annotations, making manual curation impractical. To address this gap, we introduce MAGRET, a large-scale dataset for multitarget RSVG images with cross-modal annotations. MAGRET is constructed using an automatic annotation pipeline assisted by multimodal large language models (MLLMs), substantially reducing annotation cost. Each annotation aligns natural language descriptions with multiple targets, capturing object attributes, quantities, spatial relations, and contextual details in complex scenes. We further perform rigorous test case validation and correction to ensure the reliability and structural quality of all generated descriptions. Leveraging this pipeline, we further release dataset variants with varying description lengths to enable a systematic study of language complexity in grounding performance. Alongside MAGRET, we provide comprehensive annotations, a dedicated evaluation protocol, and a strawman solution to support benchmarking. Experiments show that MAGRET establishes a scalable and effective foundation for advancing multitarget grounding research in remote sensing (RS). The dataset and code are available at https://github.com/jiehaoxue/MAGRET

Citation format

XUE, Jiehao, et al. MAGRET: A dataset for multitarget visual grounding in remote sensing images with cross-modal annotations. IEEE Geoscience and Remote Sensing Letters, 2026, 23: 1–5.