DiffArtist: Towards Structure and Appearance Controllable Image Stylization
Ruixiang Jiang, Chang Wen Chen
Abstract
Artistic styles are defined by both their structural and appearance elements. Existing neural stylization techniques primarily focus on transferring appearance-level features such as color and texture, often neglecting the equally crucial aspect of structural stylization. To address this gap, we introduce DiffArtist, the first 2D stylization method to offer fine-grained, simultaneous control over both structure and appearance style strength. This dual controllability is achieved by representing structure and appearance generation as separate diffusion processes, necessitating no further tuning or additional adapters. To properly evaluate this new capability of dual stylization, we further propose a Multimodal LLM-based stylization evaluator that aligns significantly better with human preferences than existing metrics. Extensive analysis shows that DiffArtist achieves superior style fidelity and dual-controllability compared to state-of-the-art methods. Its text-driven, training-free design and unprecedented dual controllability make it a powerful and interactive tool for various creative applications. Project homepage: https://diffusionartist.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang et al.NeurIPS 2025 · 12 citations
- Multimodal LLMs Can Reason about Aesthetics in Zero-ShotRuixiang Jiang, Chang Wen ChenACM MM 2025 · 6 citations
- StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image ReferencesBoyu He, Yunfan Ye, Chang Liu, Weishang Wu et al.CVPR 2026 · 3 citations
- CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style TransferWenbo Nie, Zixiang Li, Renshuai Tao, Bin WU et al.ICLR 2026 · 2 citations
- DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style PersonalizationAniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das et al.ICCV 2025 · 1 citation
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
Related papers
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionBowen Fu, Wei Wei, Jiaqi Tang, Jiangtao Nie et al.ICCV 2025 · 2 citations
- Pastiche Master: Exemplar-Based High-Resolution Portrait Style TransferShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 130 citations
- Dual-head Genre-instance Transformer Network for Arbitrary Style TransferMeichen Liu, Shuting He, Songnan Lin, Bihan WenACM MM 2024 · 3 citations
- DualAST: Dual Style-Learning Networks for Artistic Style TransferHaibo Chen, Lei Zhao, Zhizhong Wang, Huiming Zhang et al.CVPR 2021
