R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models
Huixuan Zhang, Xiaojun Wan
Abstract
Text-to-image models frequently fail to achieve perfect alignment with textual prompts, particularly in maintaining proper semantic binding between semantic elements in the given prompt. Existing approaches typically require costly retraining or focus on only correctly generating the attributes of entities (entityattribute binding), ignoring the cruciality of correctly generating the relations between entities (entity-relation-entity binding), resulting in unsatisfactory semantic binding performance. In this work, we propose a novel trainingfree method R-Bind that simultaneously improves both entity-attribute and entity-relationentity binding. Our method introduces three inference-time optimization losses that adjust attention maps during generation. Comprehensive evaluations across multiple datasets demonstrate our approach's effectiveness, validity, and flexibility in enhancing semantic binding without additional training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc1bdfc2-a9ff-4dd2-87ca-256f631297fdBuilds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map AlignmentRoyi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel et al.NeurIPS 2023 · 212 citations
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao et al.NeurIPS 2024 · 45 citations
- DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingXin Xie, Dong GongCVPR 2025
- RAGD: Regional-Aware Diffusion Model for Text-to-Image GenerationZhennan Chen, Yajie Li, Haofan Wang, Zhibo Chen et al.ICCV 2025 · 3 citations
- Semantic Alignment for Pose-Invariant Identity Preserving DiffusionJiwon Kim, Seonhwa Kim, Soobin Park, Eunju Cha et al.CVPR 2026
