An Image-like Diffusion Method for Human-Object Interaction Detection
Xiaofei Hui, Haoxuan Qu, Hossein Rahmani, Jun Liu
Abstract
Human-object interaction (HOI) detection often faces high levels of ambiguity and indeterminacy, as the same interaction can appear vastly different across different humanobject pairs. Additionally, the indeterminacy can be further exacerbated by issues such as occlusions and cluttered backgrounds. To handle such a challenging task, in this work, we begin with a key observation: the output of HOI detection for each human-object pair can be recast as an image. Thus, inspired by the strong image generation capabilities of image diffusion models, we propose a new framework, HOI-IDiff. In HOI-IDiff, we tackle HOI detection from a novel perspective, using an Image-like Diffusion process to generate HOI detection outputs as images. Furthermore, recognizing that our recast images differ in certain properties from natural images, we enhance our framework with a customized HOI diffusion process and a slice patchification model architecture, which are specifically tailored to generate our recast "HOI images". Extensive experiments demonstrate the efficacy of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aa80b0e-a33a-40ca-988d-2ec04a9e107bCited by top-tier papers2
- DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image GenerationZhuoling Li, Hossein Rahmani, Jiarui Zhang, Yu Xue et al.CVPR 2026 · 5 citations
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token SpacePanqi Yang, Haodong Jing, Nanning Zheng, Yongqiang MaAAAI 2026 · 2 citations
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré et al.NeurIPS 2021 · 782 citations
Related papers
- Visual Relation Diffusion for Human-Object Interaction DetectionPing Cao, Yepeng Tang, Chunjie Zhang, Xiaolong Zheng et al.ICCV 2025 · 1 citation
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction ScenariosLingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min et al.NeurIPS 2025 · 12 citations
- HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction AwarenessZihui Xue, Romy Luo, Changan Chen, Kristen GraumanNeurIPS 2024 · 29 citations
- Learning to Generate Human-Human-Object Interactions from Textual DescriptionsJeonghyeon Na, Sangwon Baik, Inhee Lee, Junyoung Lee et al.NeurIPS 2025 · 3 citations
