Sizhuang Zhang, Ying Sun, Derui Ding, Hui Yu
Abstract
High-fidelity 3-D face reconstruction is critical for enhancing personalized and immersive human–machine interaction experiences. However, existing methods struggle to capture the full spectrum of facial textures, particularly fine-scale details, such as wrinkles and pores, due to limitations in multiscale representation. To address this challenge, we propose a self-supervised multiscale hierarchical network to hierarchically model fine geometric details in multiple scales in this study. We design a global and local Markov random field loss and a detail perception loss to provide a global and local sensory field of view guidance for retaining fine-scale detail structure information of the face. In addition, we introduce a learnable Gabor-aware texture enhancement module to enhance the network’s sensitivity to fine textures. Extensive experiments show that the proposed method can reconstruct fine-scale details of the face and has superior performance to the state-of-the-art methods in terms of reconstruction accuracy and visual effect.
Citation format
ZHANG, Sizhuang, et al. Smhnet: Self-supervised multiscale hierarchical network for high fidelity 3-d face reconstruction. IEEE Transactions on Human-Machine Systems, 2026, 56(1): 114–123.