Sheng Xu, Zhongpei Sun, Hong Xu, Wei Huang, Qiqiang Chen, Hongjie He
2026IEEE Access
Abstract
Detecting small and occluded objects in complex traffic scenes remains a major challenge for intelligent transportation systems, including the difficulty in balancing high accuracy and low computational cost. To address these challenges, this paper proposes an innovative MSF-YOLOv11n model. With a modular architecture design, we propose three innovative technical frameworks. First, the Adaptive Attention-Enhanced Wavelet Module (AAEWM) performs multi-scale wavelet decomposition combined with attention-guided reconstruction, enabling the model to capture structural information such as edges and textures while suppressing noise interference. Second, the Feature-Enhanced Attention Upsampling Module (FEAUM) integrates lightweight Ghost convolution with Shuffle Attention and employs a hybrid interpolation and convolution strategy, which preserves fine-grained details in the upsampling process. Third, the Multi-Scale Fusion Lightweight Triple Attention Module (MFTAM) employs depthwise separable convolutions and a triple attention mechanism across channel, spatial and positional dimensions, effectively enhancing multi-scale feature interaction and improving small-object localization accuracy. Compared with YOLOv11n, MSF-YOLOv11n improves mAP50-95 by 3.8% and mAP50 by 3.7% on UA-DETRAC, and gains 3.6% in mAP50-95 and 3.9% in mAP50 on DAIR-V2X, all with a computational cost of only 6.5 GFLOPs and an inference speed of 100 FPS on UA-DETRAC. These results demonstrate that MSF-YOLOv11n achieves a strong balance of accuracy and efficiency for real-world traffic perception.
Citation format
XU, Sheng, et al. MSF-YOLOv11n: A multi-scale feature fusion-based efficient detector for small and occluded object detection in complex traffic scenes. IEEE Access, 2026, 14: 81515–81534.