Separate-and-Enhance: Compositional Finetuning for Text-to-Image Diffusion Models
Zhipeng Bao, Yijun Li, Krishna Kumar Singh, Yu-Xiong Wang, Martial Hebert
Abstract
Despite recent significant strides achieved by diffusion-based Text-to-Image (T2I) models, current systems are still less capable of ensuring decent compositional generation aligned with text prompts, particularly for the multi-object generation. In this work, we first show the fundamental reasons for such misalignment by identifying issues related to low attention activation and mask overlaps. Then we propose a compositional finetuning framework with two novel objectives, the Separate loss and the Enhance loss, that reduce object mask overlaps and maximize attention scores, respectively. Unlike conventional test-time adaptation methods, our model, once finetuned on critical parameters, is able to directly perform inference given an arbitrary multi-object prompt, which enhances the scalability and generalizability. Through comprehensive evaluations, our model demonstrates superior performance in image realism, text-image alignment, and adaptability, significantly surpassing established baselines. Furthermore, we show that training our model with a diverse range of concepts enables it to generalize effectively to novel concepts, exhibiting enhanced performance compared to models trained on individual concept pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b98d369f-ded7-4c5c-8ae8-25c236523abbCited by top-tier papers10
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 22 citations
- Understanding Multi-Granularity for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Seungho Lee, Minhyun Lee et al.NeurIPS 2024 · 7 citations
- Steer away from Mode Collisions: Improving Composition in Diffusion ModelsDebottam Dutta, Jianchong Chen, Rajalaxmi Rajagopalan, Yu-Lin Wei et al.ICLR 2026 · 6 citations
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationGuanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu et al.ICCV 2025 · 4 citations
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman et al.SIGGRAPH 2025 · 3 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- A-STAR: Test-time Attention Segregation and Retention for Text-to-image SynthesisAishwarya Agarwal, Srikrishna Karanam, K. J. Joseph, Apoorv Saxena et al.ICCV 2023 · 77 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingDongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang et al.NeurIPS 2024 · 75 citations
- RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion ModelsXinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu et al.NeurIPS 2024 · 22 citations
