Lune

EMNLP2025Top-tier venue

R-Bind: Unified Enhancement of Attribute and Relation Binding in Text-to-Image Diffusion Models

Huixuan Zhang, Xiaojun Wan

2025Year

Abstract

Text-to-image models frequently fail to achieve perfect alignment with textual prompts, particularly in maintaining proper semantic binding between semantic elements in the given prompt. Existing approaches typically require costly retraining or focus on only correctly generating the attributes of entities (entityattribute binding), ignoring the cruciality of correctly generating the relations between entities (entity-relation-entity binding), resulting in unsatisfactory semantic binding performance. In this work, we propose a novel trainingfree method R-Bind that simultaneously improves both entity-attribute and entity-relationentity binding. Our method introduces three inference-time optimization losses that adjust attention maps during generation. Comprehensive evaluations across multiple datasets demonstrate our approach's effectiveness, validity, and flexibility in enhancing semantic binding without additional training.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fc1bdfc2-a9ff-4dd2-87ca-256f631297fd

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines