CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning
Longchen Dai, Zixuan Shen, Zhiheng Zhou, Peipeng Yu, Zhihua Xia
Abstract
Face recognition systems store face templates for efficient matching. Once leaked, these templates pose a threat: inverting them can yield photorealistic surrogates that compromise privacy and enable impersonation. Although existing research has achieved relatively realistic face template inversion, the reconstructed facial images exhibit over-smoothed facial-part attributes (eyes, nose, mouth) and limited transferability. To address this problem, we present CLIP-FTI, a CLIP-driven fine-grained attribute conditioning framework for face template inversion. Our core idea is to use the CLIP model to obtain the semantic embeddings of facial features, in order to realize the reconstruction of specific facial feature attributes. Specifically, facial feature attribute embeddings extracted from CLIP are fused with the leaked template via a cross-modal feature interaction network and projected into the intermediate latent space of a pretrained Style- GAN. The StyleGAN generator then synthesizes face images with the same identity as the templates but with more finegrained facial feature attributes. Experiments across multiple face recognition backbones and datasets show that our reconstructions (i) achieve higher identification accuracy and attribute similarity, (ii) recover sharper component-level attribute semantics, and (iii) improve cross-model attack transferability compared to prior reconstruction attacks. To the best of our knowledge, ours is the first method to use additional information besides the face template attack to realize face template inversion and obtains SOTA results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano et al.SIGGRAPH 2022 · 501 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit DirectionsRameen Abdal, Peihao Zhu, John Femiani, Niloy J. Mitra et al.SIGGRAPH 2022 · 76 citations
Related papers
- MIRROR: Model Inversion for Deep LearningNetwork with High FidelityGuanhong Tao, Qiuling Xu, Yingqi Liu, Guangyu Shen et al.NDSS 2022
- Face Reconstruction from Facial Templates by Learning Latent Space of a Generator NetworkHatef Otroshi-Shahreza, Sébastien MarcelNeurIPS 2023 · 48 citations
- Identity-Preserving Facial Aesthetic Enhancement via Hierarchical Prompt Learning and Pivotal TuningFangli Ying, Zhihong Zhang, Liting Zhou, Cathal Gurrin et al.ACM MM 2025
- SAT3D: Image-driven Semantic Attribute Transfer in 3DZhijun Zhai, Zengmao Wang, Xiaoxiao Long, Kaixuan Zhou et al.ACM MM 2024
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu et al.CVPR 2023
