InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
Jiun Tian Hoe, Xudong Jiang, Chee Seng Chan, Yap-Peng Tan, Weipeng Hu
2024Year
20Top-tier citations
Abstract
Caption: a person is holding a bag, another person is talking on a cell phone Caption: a person is feeding a cat Generated Image Generated Image Generated Image person another person person feeding person cell phone cat another person person talking holding bag bag cell phone cat Figure 1. Generated samples of size 512x512. Stable Diffusion conditions on text caption only, while GLIGEN conditions on extra layout input. Our proposed InteractDiffusion conditions on extra interaction label and its location shown by the shaded area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesTianlong Xu, Chen Wang, Gaoyang Liu, Yang Yang et al.NeurIPS 2024 · 17 citations
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang et al.ICCV 2025 · 13 citations
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao et al.ICCV 2025 · 7 citations
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 6 citations
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong et al.ICCV 2025 · 5 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- VerbDiff: Text-Only Diffusion Models with Enhanced Interaction AwarenessSeungJu Cha, Kwanyoung Lee, Ye-Chan Kim, Hyunwoo Oh et al.CVPR 2025
- Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsShuangqi Li, Hieu Le, Jingyi Xu, Mathieu SalzmannICLR 2025
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio ConditionsZhenzhi Wang, Jiaqi Yang, Jianwen Jiang, Chao Liang et al.ICLR 2026 · 19 citations
- LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image GenerationGuangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi et al.CVPR 2023
- Diffusion Model is Effectively Its Own TeacherXinyin Ma, Runpeng Yu, Songhua Liu, Gongfan Fang et al.CVPR 2025
