Cybersecurity and Information SystemsAdvanced Data Processing TechniquesNeural Networks and Reservoir Computing

Kozynets A, Demkiv L

2026.3.30ARTIFICIAL INTELLIGENCE

DOI: 10.15407/jai2026.01.059

Abstract

The article considers the problem of effective knowledge transfer between deep convolutional neural networks of different capacities for their further deployment on hardware platforms with limited computing resources. The main problem of standard knowledge distillation protocols is the occurrence of "gradient shock" during the initialization of the student model on specific data sets, which leads to the destruction of the feature space and the loss of final accuracy. To overcome this limitation, the Two-Stage Distillation algorithm was developed and implemented. The proposed approach divides the learning process into the classifier stabilization phase and the deep distillation phase. Experimental research was conducted on the ResNet and VGG architectural families. The results obtained confirm that the use of the proposed algorithm allows to increase the accuracy of compact models by 1.5–2.6% compared to standard training. In addition, the work investigated and experimentally confirmed the phenomenon of "distillation recovery" - the ability of the algorithm to restore the accuracy of the model after aggressive structural pruning. It is proven that the use of "soft goals" of the teacher in a narrowed search space allows the sparse ResNet18 model to achieve an accuracy of 77.8%, which exceeds the basic full-size student model

Citation format

A, Kozynets; L, Demkiv. Research on two-stage knowledge distillation and hybrid compression of convolutional neural networks to reduce their computational complexity. ARTIFICIAL INTELLIGENCE, 2026, 31(AI.2026.31(1)): 59–69.