StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models
Zhizhong Wang, Lei Zhao, Wei Xing
Abstract
Content and style (C-S) disentanglement is a fundamental problem and critical challenge of style transfer. Existing approaches based on explicit definitions (e.g., Gram matrix) or implicit learning (e.g., GANs) are neither interpretable nor easy to control, resulting in entangled representations and less satisfying results. In this paper, we propose a new C-S disentangled framework for style transfer without using previous assumptions. The key insight is to explicitly extract the content information and implicitly learn the complementary style information, yielding interpretable and controllable C-S disentanglement and style transfer. A simple yet effective CLIP-based style disentanglement loss coordinated with a style reconstruction prior is introduced to disentangle C-S in the CLIP image space. By further leveraging the powerful style removal and generative ability of diffusion models, our framework achieves superior results than state of the art and flexible C-S disentanglement and trade-off control. Our work provides new insights into the C-S disentanglement in style transfer and demonstrates the potential of diffusion models for learning well-disentangled C-S characteristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6166cc97-c66b-4e9c-8a77-7acb7869bc63Cited by top-tier papers74
- Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image EditingHaonan Lin, Yan Chen, Jiahao Wang, Wenbin An et al.NeurIPS 2024 · 46 citations
- How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang et al.NeurIPS 2024 · 42 citations
- BELM: Bidirectional Explicit Linear Multi-step Sampler for Exact Inversion in Diffusion ModelsFangyikang Wang, Hubery Yin, Yuejiang Dong, Huminhao Zhu et al.NeurIPS 2024 · 37 citations
- AsyncDiff: Parallelizing Diffusion Models by Asynchronous DenoisingZigeng Chen, Xinyin Ma, Gongfan Fang, Zhenxiong Tan et al.NeurIPS 2024 · 33 citations
- Diffusion Handles Enabling 3D Edits for Diffusion Models by Lifting Activations to 3DKarran Pandey, Paul Guerrero, Matheus Gadelha, Yannick Hold-Geoffroy et al.CVPR 2024 · 19 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- CLIPstyler: Image Style Transfer with a Single Text ConditionGihyun Kwon, Jong Chul YeCVPR 2022 · 224 citations
- Diffusion-based Image Translation using disentangled style and content representationGihyun Kwon, Jong Chul YeICLR 2023 · 46 citations
- StyleDistillation: A New Insight of Image Style Enables Personalized Aesthetic ManipulationYuxin Wang, Xiaoyu Geng, Yuke Li, Zheng WangICML 2026
- Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion ModelsLin Zhu, Xinbing Wang, Chenghu Zhou, Qinying Gu et al.ICLR 2025
- Disentangled Learning with Synthetic Parallel Data for Text Style TransferJingxuan Han, Quan Wang, Zikang Guo, Benfeng Xu et al.ACL 2024 · 4 citations
