ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors
Jingwen Chen, Yingwei Pan, Ting Yao, Tao Mei
Abstract
Recently, the multimedia community has witnessed the rise of diffusion models trained on large-scale multi-modal data for visual content creation, particularly in the field of text-to-image generation. In this paper, we propose a new task for "stylizing'' text-to-image models, namely text-driven stylized image generation, that further enhances editability in content creation. Given input text prompt and style image, this task aims to produce stylized images which are both semantically relevant to input text prompt and meanwhile aligned with the style image in style. To achieve this, we present a new diffusion model (ControlStyle) via upgrading a pre-trained text-to-image model with a trainable modulation network enabling more conditions of text prompts and style images. Moreover, diffusion style and content regularizations are simultaneously introduced to facilitate the learning of this modulation network with these diffusion priors, pursuing high-quality stylized text-to-image generation. Extensive experiments demonstrate the effectiveness of our ControlStyle in producing more visually pleasing and artistic results, surpassing a simple combination of text-to-image model and conventional style transfer techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e810a928-d60b-40d2-ac0c-27131f90e590Cited by top-tier papers16
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image GenerationMushui Liu, Yuhang Ma, Zhen Yang, Jun Dan et al.AAAI 2025 · 36 citations
- Block and Detail: Scaffolding Sketch-to-Image GenerationVishnu Sarukkai, Lu Yuan, Mia Tang, Maneesh Agrawala et al.UIST 2024 · 23 citations
- Boosting Diffusion Models with Moving Average Sampling in Frequency DomainYurui Qian, Qi Cai, Yingwei Pan, Yehao Li et al.CVPR 2024 · 22 citations
- SCott: Accelerating Diffusion Models with Stochastic Consistency DistillationHongjian Liu, Qingsong Xie, Tianxiang Ye, Zhijie Deng et al.AAAI 2025 · 17 citations
- SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*Rui Zhu, Yingwei Pan, Yehao Li, Ting Yao et al.CVPR 2024 · 15 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- StyleDistillation: A New Insight of Image Style Enables Personalized Aesthetic ManipulationYuxin Wang, Xiaoyu Geng, Yuke Li, Zheng WangICML 2026
- Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion ModelsLin Zhu, Xinbing Wang, Chenghu Zhou, Qinying Gu et al.ICLR 2025
- Style Aligned Image Generation via Shared AttentionAmir Hertz, Andrey Voynov, Shlomi Fruchter, Daniel Cohen-OrCVPR 2024
- Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationXiang Gao, Zhengbo Xu, Junhan Zhao, Jiaying LiuAAAI 2024 · 23 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
