Deformable One-Shot Face Stylization via DINO Semantic Guidance
Yang Zhou, Zichong Chen, Hui Huang
Abstract
This paper addresses the complex issue of one-shot face stylization, focusing on the simultaneous consideration of appearance and structure, where previous methods have fallen short. We explore deformation-aware face stylization that diverges from traditional single-image style reference, opting for a real-style image pair instead. The cornerstone of our method is the utilization of a self-supervised vision transformer, specifically DINO-ViT, to establish a robust and consistent facial structure representation across both real and style domains. Our stylization process begins by adapting the StyleGAN generator to be deformation-aware through the integration of spatial transformers (STN). We then introduce two innovative constraints for generator fine-tuning under the guidance of DINO semantics: i) a directional deformation loss that regulates directional vectors in DINO space, and ii) a relative structural consistency constraint based on DINO token self-similarities, ensuring diverse generation. Additionally, style-mixing is employed to align the color generation with the reference, minimizing inconsistent correspondences. This framework delivers enhanced deformability for general one-shot face stylization, achieving notable efficiency with a fine-tuning duration of approximately 10 minutes. Extensive qualitative and quantitative comparisons demonstrate our superiority over state-of-the-art one-shot face stylization methods. Code is available at https://github.com/zichongc/DoesFS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2bcbae8-b861-4279-8e6a-dba29ef22484Cited by top-tier papers5
- ZePo: Zero-Shot Portrait Stylization with Faster SamplingJin Liu, Huaibo Huang, Jie Cao, Ran HeACM MM 2024 · 6 citations
- Neural-Driven Image EditingPengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao et al.NeurIPS 2025 · 5 citations
- DiffArtist: Towards Structure and Appearance Controllable Image StylizationRuixiang Jiang, Chang Wen ChenACM MM 2025 · 4 citations
- Latent Space ImagingMatheus Souza, Yidan Zheng, Kaizhang Kang, Yogeshwar Nath Mishra et al.CVPR 2025
- FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake VideosZhaolun Li, Jichang Li, Yinqi Cai, Junye Chen et al.ICCV 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
Related papers
- Towards Diverse and Faithful One-shot Adaption of Generative Adversarial NetworksYabo Zhang, Mingshuai Yao, Yuxiang Wei, Zhilong Ji et al.NeurIPS 2022 · 30 citations
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan et al.AAAI 2023 · 135 citations
- Stylos: Multi-View 3D Stylization with Single-Forward Gaussian SplattingHanzhou Liu, Jia Huang, Mi Lu, Srikanth Saripalli et al.ICLR 2026 · 4 citations
- Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentQing Chang, Yao-Xiang Ding, Kun ZhouAAAI 2025 · 3 citations
- DeformToon3d: Deformable Neural Radiance Fields for 3D ToonificationJunzhe Zhang, Yushi Lan, Shuai Yang, Fangzhou Hong et al.ICCV 2023 · 16 citations
