Relation Rectification in Diffusion Model
Yinwei Wu, Xingyi Yang, Xinchao Wang
Abstract
Despite their exceptional generative abilities, large T2I diffusion models, much like skilled but careless artists, often struggle with accurately depicting visual relationships between objects. This issue, as we uncover through careful analysis, arises from a misaligned text encoder that struggles to interpret specific relationships and differentiate the logical order of associated objects. To resolve this, we in-troduce a novel task termed Relation Rectification, aiming to refine the model to accurately represent a given relationship it initially fails to generate. To address this, we propose an innovative solution utilizing a Heterogeneous Graph Convolutional Network (HGCN). It models the di-rectional relationships between relation terms and corre-sponding objects within the input prompts. Specifically, we optimize the HGCN on a pair of prompts with identical relational words but reversed object orders, supplemented by a few reference images. The lightweight HGCN adjusts the text embeddings generated by the text encoder, ensuring accurate reflection of the textual relation in the em-bedding space. Crucially, our method retains the parameters of the text encoder and diffusion model, preserving the model's robust performance on unrelated descriptions. We validated our approach on a newly curated dataset of di-verse relational data, demonstrating both quantitative and qualitative enhancements in generating images with precise visual relations. Project page: https://wuyinwei-hah.github.io/rrnet.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao et al.NeurIPS 2024 · 45 citations
- Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate DecompositionSina Malakouti, Adriana KovashkaNeurIPS 2025 · 5 citations
- IFAdapter: Instance Feature Control for Grounded Text-to-Image GenerationYinwei Wu, Xianpan Zhou, Bing Ma, Xuefeng Su et al.ICCV 2025 · 3 citations
- SUV: Suppressing Undesired Video Content via Semantic Modulation Based on Text EmbeddingsXiang Lv, Mingwen Shao, Lingzhuang Meng, Chang Liu et al.ICCV 2025 · 2 citations
- Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationGuibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge et al.ICCV 2025 · 1 citation
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- DreamRelation: Bridging Customization and Relation GenerationQingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai et al.CVPR 2025
- HGCN: A Heterogeneous Graph Convolutional Network-Based Deep Learning Model Toward Collective ClassificationZhihua Zhu, Xinxin Fan, Xiaokai Chu, Jingping BiKDD 2020 · 46 citations
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing et al.ACM MM 2022 · 59 citations
- Grounded Image Text Matching with Mismatched Relation ReasoningYu Wu, Yana Wei, Haozhe Wang, Yongfei Liu et al.ICCV 2023 · 14 citations
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 70 citations
