Relation Rectification in Diffusion Model
Yinwei Wu, Xingyi Yang, Xinchao Wang
摘要
Despite their exceptional generative abilities, large T2I diffusion models, much like skilled but careless artists, often struggle with accurately depicting visual relationships between objects. This issue, as we uncover through careful analysis, arises from a misaligned text encoder that struggles to interpret specific relationships and differentiate the logical order of associated objects. To resolve this, we in-troduce a novel task termed Relation Rectification, aiming to refine the model to accurately represent a given relationship it initially fails to generate. To address this, we propose an innovative solution utilizing a Heterogeneous Graph Convolutional Network (HGCN). It models the di-rectional relationships between relation terms and corre-sponding objects within the input prompts. Specifically, we optimize the HGCN on a pair of prompts with identical relational words but reversed object orders, supplemented by a few reference images. The lightweight HGCN adjusts the text embeddings generated by the text encoder, ensuring accurate reflection of the textual relation in the em-bedding space. Crucially, our method retains the parameters of the text encoder and diffusion model, preserving the model's robust performance on unrelated descriptions. We validated our approach on a newly curated dataset of di-verse relational data, demonstrating both quantitative and qualitative enhancements in generating images with precise visual relations. Project page: https://wuyinwei-hah.github.io/rrnet.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Token Merging for Training-Free Semantic Binding in Text-to-Image SynthesisTaihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao 等NeurIPS 2024 · 被引用 45 次
- Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate DecompositionSina Malakouti, Adriana KovashkaNeurIPS 2025 · 被引用 5 次
- IFAdapter: Instance Feature Control for Grounded Text-to-Image GenerationYinwei Wu, Xianpan Zhou, Bing Ma, Xuefeng Su 等ICCV 2025 · 被引用 3 次
- SUV: Suppressing Undesired Video Content via Semantic Modulation Based on Text EmbeddingsXiang Lv, Mingwen Shao, Lingzhuang Meng, Chang Liu 等ICCV 2025 · 被引用 2 次
- Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationGuibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- DreamRelation: Bridging Customization and Relation GenerationQingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai 等CVPR 2025
- HGCN: A Heterogeneous Graph Convolutional Network-Based Deep Learning Model Toward Collective ClassificationZhihua Zhu, Xinxin Fan, Xiaokai Chu, Jingping BiKDD 2020 · 被引用 46 次
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing 等ACM MM 2022 · 被引用 59 次
- Grounded Image Text Matching with Mismatched Relation ReasoningYu Wu, Yana Wei, Haozhe Wang, Yongfei Liu 等ICCV 2023 · 被引用 14 次
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 被引用 70 次
