Face recognition and analysisAdvanced Neural Network ApplicationsBiometric Identification and Security

Van-Khoa Pham, Manh-Dung Do, Trung-Nghia Dang, L. Tran

2026.3.25EAI Endorsed Transactions on AI and Robotics

DOI: 10.4108/airo.10070

tlooto Summary

Results support the practicality of explicitly separating software-stage quantization analysis from hardware-stage realization for real-time, low-power facial landmark detection on FPGA-based edge platforms.

Abstract

Facial landmark detection is a key component of always-on edge vision systems, but practical deployment requires balancing localization accuracy, model size, throughput, and power consumption. This study proposes a two-stage, hardware-aware framework for MobileNetV2-based facial landmark detection. In Stage I, a lightweight detector is developed in PyTorch and evaluated in FP32 and INT8 using post-training quantization (PTQ) and quantization-aware training (QAT). In Stage II, the quantized model is realized on the AMD/Xilinx Kria KV260 FPGA and assessed in terms of real-time throughput, power consumption, and hardware resource utilization. INT8 quantization reduces the model size from 6.59 MB to 1.65 MB, and QAT retains accuracy more effectively than PTQ (91.74% vs. 90.94%) relative to the FP32 baseline (92.42%). The hardware implementation achieves approximately 30 FPS at approximately 3 W while using 14.8% of LUTs, 8.1% of FFs, 16.3% of BRAM, and 4.5% of DSPs. Among the evaluated platforms, the KV260 delivers the highest measured energy efficiency, whereas the RTX 4060 delivers the highest throughput. Within the evaluated setup, these results support the practicality of explicitly separating software-stage quantization analysis from hardware-stage realization for real-time, low-power facial landmark detection on FPGA-based edge platforms.

Citation format

PHAM, Van-Khoa, et al. Hardware-aware INT8 quantization and FPGA deployment of mobilenetv2 for real-time facial landmark detection. EAI Endorsed Transactions on AI and Robotics, 2026, 5.