Yeeun Choi, H. Jeong, Chaemin Yoo, Gwanghee Lee, Kyoungson Jhang
Abstract
Facial landmark detection is a task that involves estimating landmark points such as eyes, nose, and mouth in facial images, and is utilized in various applications including facial recognition, emotion analysis, expression recognition, and person identification. Among implementation methods, coordinate regression faces a significant challenge with notably lower prediction accuracy in specific areas such as the jaw and below the ears. We propose a selective enhancement of soft-argmax estimation (SESAME) as a technique to address these limitations. SESAME consists of three stages. In the preliminary step, we train a randomly initialized facial landmark detection model using the WFLW facial landmark dataset to identify areas of poor performance. During the pre-training step, we focus on enhancing representation learning by applying intensive masking to poorly performing landmark regions in random facial input images from WFLW, followed by reconstruction. In the fine-tuning step, we further train the enhanced model using the WFLW dataset. Experimental results show significant improvement in the prediction accuracy of traditionally challenging areas, particularly in jawline regions where the normalized mean error (NME) substantially decreased from 6.547% to 6.184%. The overall average NME across all regions improved from 4.397% to 4.351%, demonstrating enhanced overall performance.
Citation format
CHOI, Yeeun, et al. SESAME: Selective enhancement of soft-argmax estimation for challenging facial landmark detection via adaptive masking- based representation learning. Journal of Computing Science and Engineering, 2025, 19: 72–80.