Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
Dongmin Park, Sebin Kim, Taehong Moon, Minkyu Kim, Kangwook Lee, Jaewoong Cho
Abstract
State-of-the-art text-to-image (T2I) diffusion models often struggle to generate rare compositions of concepts, e.g., objects with unusual attributes. In this paper, we show that the compositional generation power of diffusion models on such rare concepts can be significantly enhanced by the Large Language Model (LLM) guidance. We start with empirical and theoretical analysis, demonstrating that exposing frequent concepts relevant to the target rare concepts during the diffusion sampling process yields more accurate concept composition. Based on this, we propose a training-free approach, R2F, that plans and executes the overall rare-to-frequent concept guidance throughout the diffusion inference by leveraging the abundant semantic knowledge in LLMs. Our framework is flexible across any pre-trained diffusion models and LLMs, and can be seamlessly integrated with the region-guided diffusion approaches. Extensive experiments on three datasets, including our newly proposed benchmark, RareBench, containing various prompts with rare compositions of concepts, R2F significantly surpasses existing models including SD3.0 and FLUX by up to 28.1%p in T2I alignment. Code is available at https://github.com/krafton-ai/Rare-to-Frequent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Image Generation from Contextually-Contradictory PromptsSaar Huberman, Or Patashnik, Omer Dahary, Ron Mokady et al.CVPR 2026 · 11 citations
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-stepZheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang et al.NeurIPS 2025 · 9 citations
- AR-RAG: Autoregressive Retrieval Augmentation for Image GenerationJingyuan Qi, Zhiyang Xu, Qifan Wang, Lifu HuangNeurIPS 2025 · 7 citations
- Rare Text Semantics Were Always There in Your Diffusion TransformerSeil Kang, Woojung Han, Dayun Ju, Seong Jae HwangNeurIPS 2025 · 6 citations
- Steer away from Mode Collisions: Improving Composition in Diffusion ModelsDebottam Dutta, Jianchong Chen, Rajalaxmi Rajagopalan, Yu-Lin Wei et al.ICLR 2026 · 6 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion GenerationKwanyoung Lee, SeungJu Cha, Yebin Ahn, Hyunwoo Oh et al.CVPR 2026
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image GenerationMushui Liu, Yuhang Ma, Zhen Yang, Jun Dan et al.AAAI 2025 · 36 citations
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- Re-Imagen: Retrieval-Augmented Text-to-Image GeneratorWenhu Chen, Hexiang Hu, Chitwan Saharia, William W. CohenICLR 2023 · 44 citations
- TF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionShilin Lu, Yanzhu Liu, Adams Wai-Kin KongICCV 2023 · 214 citations
