EngineeringComputer Science

Yanyu Wan, Junyan Chen, Lei Xiao, Biao Tian, Wenzhen Wu

2026.1.1IEEE TRANSACTIONS ON AEROSPACE AND ELECTRONIC SYSTEMS

DOI: 10.1109/taes.2026.3651663

Abstract

The accurate classification of low, slow, and small targets in low-altitude airspace is a critical challenge for surveillance and safety systems. While radar provides essential data, methods relying on a single information source, such as micro-Doppler signatures or kinematic tracks, are often unreliable due to low radar cross-sections and complex clutter environments. Consequently, multimodal fusion has emerged as a promising direction. However, the performance of conventional fusion algorithms is often compromised by their static fusion strategies, where fixed weights assigned during training cannot adapt to real-world scenarios in which the quality of a data stream may degrade. To address these limitations, this article proposes the multimodal adaptive feature fusion network (MAFF-Net), a multilevel dual-branch architecture designed to overcome the vulnerabilities of both single-source methods and static fusion. At each hierarchical level, a semantic bridge module first aligns heterogeneous features from different modalities using contrastive learning. Crucially, an adaptive fusion module then dynamically arbitrates the contribution of each information stream, enabling the network to intelligently rely on higher quality data and mitigate the effects of signal distortion or loss. Evaluated on a comprehensive dataset collected from five distinct geographical sites, MAFF-Net achieves 91.01% accuracy, outperforming other advanced fusion methods by up to 2.0 percentage points. The model’s robustness and practical utility are further validated through extensive cross-site testing and a successful in-the-field deployment.

Citation format

WAN, Yanyu, et al. MAFF-Net: A deep fusion framework for multimodal recognition of LSS targets. IEEE TRANSACTIONS ON AEROSPACE AND ELECTRONIC SYSTEMS, 2026, 62: 4375–4392.