CSGO: Content-Style Composition in Text-to-Image Generation
Peng Xing, Haofan Wang, Yanpeng Sun, Qixun Wang, Xu Bai, Hao Ai, Jen-Yuan Huang, Zechao Li
摘要
The advancement of image style transfer has been fundamentally constrained by the absence of large-scale, high-quality datasets with explicit content-style-stylized supervision. Existing methods predominantly adopt training-free paradigms (e.g., image inversion), which limit controllability and generalization due to the lack of structured triplet data. To bridge this gap, we design a scalable and automated pipeline that constructs and purifies high-fidelity content-style-stylized image triplets. Leveraging this pipeline, we introduce IMAGStyle-the first large-scale dataset of its kind, containing 210K diverse and precisely aligned triplets for style transfer research. Empowered by IMAGStyle, we propose CSGO, a unified, endto-end trainable framework that decouples content and style representations via independent feature injection. CSGO jointly supports image-driven style transfer, text-driven stylized generation, and text-editing-driven stylized synthesis within a single architecture. Extensive experiments show that CSGO achieves state-ofthe-art controllability and fidelity, demonstrating the critical role of structured synthetic data in unlocking robust and generalizable style transfer. Source code:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang 等NeurIPS 2025 · 被引用 12 次
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image GenerationZhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao 等CVPR 2026 · 被引用 7 次
- AlignedGen: Aligning Style Across Generated ImagesJiexuan Zhang, Yiheng Du, Qian Wang, Weiqi Li 等NeurIPS 2025 · 被引用 6 次
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang 等ICCV 2025 · 被引用 6 次
- SplitFlux: Learning to Decouple Content and Style from a Single ImageYitong Yang, Yinglin Wang, Changshuo Wang, Yongjun Zhang 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Disentangled Learning with Synthetic Parallel Data for Text Style TransferJingxuan Han, Quan Wang, Zikang Guo, Benfeng Xu 等ACL 2024 · 被引用 4 次
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 被引用 45 次
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng 等ACM MM 2022 · 被引用 8 次
- SigStyle: Signature Style Transfer via Personalized Text-to-Image ModelsYe Wang, Tongyuan Bai, Xuping Xie, Zili Yi 等AAAI 2025 · 被引用 5 次
- DreamStyle: A Unified Framework for Video StylizationMengtian Li, Jinshu Chen, Songtao Zhao, Wanquan Feng 等CVPR 2026 · 被引用 4 次
