Kunlong Zhao, Dawei Zhao, Liang Xiao, Yiming Nie, Yulong Huang, Yonggang Zhang
2026.3.1IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY
Abstract
Supervised learning-based visual object tracking methods rely heavily on manually annotated video data. Semi-supervised learning-based visual object trackers offer a promising alternative by balancing tracking performance with annotation costs. However, the frame-to-frame dependency and the paired input characteristic of visual object trackers make existing semi-supervised learning methods unsuitable. To solve these problems, a novel semi-supervised learning framework is proposed to train a tracker, which significantly reduces annotation dependency while maintaining strong performance. It adopts an iterative “track-then-train” paradigm tailored to the frame-to-frame dependency characteristic of tracker training samples. In the “train” step, a Dual-Sample-Teacher training method is proposed to better leverage the paired input characteristic and special training samples for semi-supervised learning. In the “track” step, a Track Compensation module—comprising Object Integrity Prediction and Intersection over Union (IoU) Prediction modules—is introduced to enhance both the quantity and quality of pseudo-labels. To further improve performance, an Iterative Training strategy is presented. Experimental results demonstrate that on the GOT-10k dataset, the baseline tracker trained using the proposed semi-supervised approach—leveraging only 1% of the annotated labels—achieves a 17.6% improvement in Average Overlap (AO) over supervised learning with equivalent label coverage.
Citation format
ZHAO, Kunlong, et al. IDSTT: Iterative dual-sample-teacher for semi-supervised visual object tracking. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2026, 36: 3638–3651.