Wankang Zeng, Zhiyang Lu, Zhu Wang, Ming Cheng, Siqi Shen, Changwang Zhang, Chenglu Wen

2026IEEE TRANSACTIONS ON MULTIMEDIA

DOI: 10.1109/tmm.2026.3699036

Abstract

Current unsupervised domain adaptive semantic segmentation (UDASS) approaches have achieved considerable success in mitigating the domain gap between synthetic and real-world data, thereby facilitating effective knowledge transfer from source to target domains. However, these methods exhibit diminished performance under the significant viewpoint discrepancies between car-mounted and drone-mounted camera setups. While some of these approaches implicitly leverage the discrepancies of text embeddings from different viewpoints to account for these variations, their language descriptions often fail to capture the distinctiveness of various categories from diverse perspectives. To address this challenge, we propose LDCView—a novel Language Descriptions guided Cross-View Domain Adaptive Semantic Segmentation framework. Specifically, our method leverages carefully designed prompts to guide a multimodal large language model in generating view-specific language descriptions for each semantic class. These descriptions are subsequently used to derive view-specific class prototypes. Furthermore, we introduce a text-image cross-attention decoder that facilitates the learning of view-specific semantic representations across both source and target domains. We also propose Cross-Encoder Attention Fusion (CEAF), a novel module that injects view-invariant and domain-agnostic semantic priors from a pre-trained CLIP vision encoder into the segmentation network via a cross-attention mechanism. Quantitative and qualitative experimental results on three challenging cross-view UDASS benchmarks demonstrate the effectiveness and superiority of LDCView over state-of-the-art methods.

Citation format

ZENG, Wankang, et al. Bridging cross-view domain gaps in semantic segmentation with language-guided representations. IEEE TRANSACTIONS ON MULTIMEDIA, 2026.