CSGO: Content-Style Composition in Text-to-Image Generation
Peng Xing, Haofan Wang, Yanpeng Sun, Qixun Wang, Xu Bai, Hao Ai, Jen-Yuan Huang, Zechao Li
Abstract
The advancement of image style transfer has been fundamentally constrained by the absence of large-scale, high-quality datasets with explicit content-style-stylized supervision. Existing methods predominantly adopt training-free paradigms (e.g., image inversion), which limit controllability and generalization due to the lack of structured triplet data. To bridge this gap, we design a scalable and automated pipeline that constructs and purifies high-fidelity content-style-stylized image triplets. Leveraging this pipeline, we introduce IMAGStyle-the first large-scale dataset of its kind, containing 210K diverse and precisely aligned triplets for style transfer research. Empowered by IMAGStyle, we propose CSGO, a unified, endto-end trainable framework that decouples content and style representations via independent feature injection. CSGO jointly supports image-driven style transfer, text-driven stylized generation, and text-editing-driven stylized synthesis within a single architecture. Extensive experiments show that CSGO achieves state-ofthe-art controllability and fidelity, demonstrating the critical role of structured synthetic data in unlocking robust and generalizable style transfer. Source code:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5d89d7d-6502-40fe-bb6e-f3b27a0c5ad6Cited by top-tier papers32
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang et al.NeurIPS 2025 · 12 citations
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image GenerationZhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao et al.CVPR 2026 · 7 citations
- AlignedGen: Aligning Style Across Generated ImagesJiexuan Zhang, Yiheng Du, Qian Wang, Weiqi Li et al.NeurIPS 2025 · 6 citations
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang et al.ICCV 2025 · 6 citations
- SplitFlux: Learning to Decouple Content and Style from a Single ImageYitong Yang, Yinglin Wang, Changshuo Wang, Yongjun Zhang et al.CVPR 2026 · 5 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Disentangled Learning with Synthetic Parallel Data for Text Style TransferJingxuan Han, Quan Wang, Zikang Guo, Benfeng Xu et al.ACL 2024 · 4 citations
- ControlStyle: Text-Driven Stylized Image Generation Using Diffusion PriorsJingwen Chen, Yingwei Pan, Ting Yao, Tao MeiACM MM 2023 · 45 citations
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng et al.ACM MM 2022 · 8 citations
- SigStyle: Signature Style Transfer via Personalized Text-to-Image ModelsYe Wang, Tongyuan Bai, Xuping Xie, Zili Yi et al.AAAI 2025 · 5 citations
- DreamStyle: A Unified Framework for Video StylizationMengtian Li, Jinshu Chen, Songtao Zhao, Wanquan Feng et al.CVPR 2026 · 4 citations
