Deformable One-Shot Face Stylization via DINO Semantic Guidance
Yang Zhou, Zichong Chen, Hui Huang
摘要
This paper addresses the complex issue of one-shot face stylization, focusing on the simultaneous consideration of appearance and structure, where previous methods have fallen short. We explore deformation-aware face stylization that diverges from traditional single-image style reference, opting for a real-style image pair instead. The cornerstone of our method is the utilization of a self-supervised vision transformer, specifically DINO-ViT, to establish a robust and consistent facial structure representation across both real and style domains. Our stylization process begins by adapting the StyleGAN generator to be deformation-aware through the integration of spatial transformers (STN). We then introduce two innovative constraints for generator fine-tuning under the guidance of DINO semantics: i) a directional deformation loss that regulates directional vectors in DINO space, and ii) a relative structural consistency constraint based on DINO token self-similarities, ensuring diverse generation. Additionally, style-mixing is employed to align the color generation with the reference, minimizing inconsistent correspondences. This framework delivers enhanced deformability for general one-shot face stylization, achieving notable efficiency with a fine-tuning duration of approximately 10 minutes. Extensive qualitative and quantitative comparisons demonstrate our superiority over state-of-the-art one-shot face stylization methods. Code is available at https://github.com/zichongc/DoesFS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ZePo: Zero-Shot Portrait Stylization with Faster SamplingJin Liu, Huaibo Huang, Jie Cao, Ran HeACM MM 2024 · 被引用 6 次
- Neural-Driven Image EditingPengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao 等NeurIPS 2025 · 被引用 5 次
- DiffArtist: Towards Structure and Appearance Controllable Image StylizationRuixiang Jiang, Chang Wen ChenACM MM 2025 · 被引用 4 次
- Latent Space ImagingMatheus Souza, Yidan Zheng, Kaizhang Kang, Yogeshwar Nath Mishra 等CVPR 2025
- FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake VideosZhaolun Li, Jichang Li, Yinqi Cai, Junye Chen 等ICCV 2025
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine 等NeurIPS 2020 · 被引用 2,345 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
相关 Paper
- Towards Diverse and Faithful One-shot Adaption of Generative Adversarial NetworksYabo Zhang, Mingshuai Yao, Yuxiang Wei, Zhilong Ji 等NeurIPS 2022 · 被引用 30 次
- StyleTalk: One-Shot Talking Head Generation with Controllable Speaking StylesYifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan 等AAAI 2023 · 被引用 135 次
- Stylos: Multi-View 3D Stylization with Single-Forward Gaussian SplattingHanzhou Liu, Jia Huang, Mi Lu, Srikanth Saripalli 等ICLR 2026 · 被引用 4 次
- Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentQing Chang, Yao-Xiang Ding, Kun ZhouAAAI 2025 · 被引用 3 次
- DeformToon3d: Deformable Neural Radiance Fields for 3D ToonificationJunzhe Zhang, Yushi Lan, Shuai Yang, Fangzhou Hong 等ICCV 2023 · 被引用 16 次
