DiffArtist: Towards Structure and Appearance Controllable Image Stylization
Ruixiang Jiang, Chang Wen Chen
摘要
Artistic styles are defined by both their structural and appearance elements. Existing neural stylization techniques primarily focus on transferring appearance-level features such as color and texture, often neglecting the equally crucial aspect of structural stylization. To address this gap, we introduce DiffArtist, the first 2D stylization method to offer fine-grained, simultaneous control over both structure and appearance style strength. This dual controllability is achieved by representing structure and appearance generation as separate diffusion processes, necessitating no further tuning or additional adapters. To properly evaluate this new capability of dual stylization, we further propose a Multimodal LLM-based stylization evaluator that aligns significantly better with human preferences than existing metrics. Extensive analysis shows that DiffArtist achieves superior style fidelity and dual-controllability compared to state-of-the-art methods. Its text-driven, training-free design and unprecedented dual controllability make it a powerful and interactive tool for various creative applications. Project homepage: https://diffusionartist.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 等NeurIPS 2025 · 被引用 12 次
- Multimodal LLMs Can Reason about Aesthetics in Zero-ShotRuixiang Jiang, Chang Wen ChenACM MM 2025 · 被引用 6 次
- StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image ReferencesBoyu He, Yunfan Ye, Chang Liu, Weishang Wu 等CVPR 2026 · 被引用 3 次
- CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style TransferWenbo Nie, Zixiang Li, Renshuai Tao, Bin WU 等ICLR 2026 · 被引用 2 次
- DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style PersonalizationAniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana 等NeurIPS 2023 · 被引用 1,192 次
相关 Paper
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionBowen Fu, Wei Wei, Jiaqi Tang, Jiangtao Nie 等ICCV 2025 · 被引用 2 次
- Pastiche Master: Exemplar-Based High-Resolution Portrait Style TransferShuai Yang, Liming Jiang, Ziwei Liu, Chen Change LoyCVPR 2022 · 被引用 130 次
- Dual-head Genre-instance Transformer Network for Arbitrary Style TransferMeichen Liu, Shuting He, Songnan Lin, Bihan WenACM MM 2024 · 被引用 3 次
- DualAST: Dual Style-Learning Networks for Artistic Style TransferHaibo Chen, Lei Zhao, Zhizhong Wang, Huiming Zhang 等CVPR 2021
