Toward Open-Set Human Object Interaction Detection
Mingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji
Abstract
This work is oriented toward the task of open-set Human Object Interaction (HOI) detection. The challenge lies in identifying completely new, out-of-domain relationships, as opposed to in-domain ones which have seen improvements in zero-shot HOI detection. To address this challenge, we introduce a simple Disentangled HOI Detection (DHD) model for detecting novel relationships by integrating an open-set object detector with a Visual Language Model (VLM). We utilize a disentangled image-text contrastive learning metric for training and connect the bottom-up visual features to text embeddings through lightweight unary and pair-wise adapters. Our model can benefit from the open-set object detector and the VLM to detect novel action categories and combine actions with novel object categories. We further present the VG-HOI dataset, a comprehensive benchmark with over 17k HOI relationships for open-set scenarios. Experimental results show that our model can detect unknown action classes and combine unknown object classes. Furthermore, it can generalize to over 17k HOI classes while being trained on just 600 HOI classes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language ModelsMingrui Wu, Jiayi Ji, Oucheng Huang, Jiale Li et al.ICML 2024 · 32 citations
- MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsDongpan Chen, Dehui Kong, Jinghua Li, Baocai YinAAAI 2025 · 5 citations
- Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction DetectionFrancesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan et al.ACM MM 2025 · 2 citations
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionChen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu et al.ACM MM 2025 · 2 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Detecting Human-Object Interactions via Functional GeneralizationAnkan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, Rama ChellappaAAAI 2020 · 131 citations
Related papers
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationQinqian Lei, Bo Wang, Robby T. TanICCV 2025 · 4 citations
- End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationMingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin et al.AAAI 2023 · 64 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction RecognitionShiyu Xuan, Dongkai Wang, Zechao Li, Jinhui TangICLR 2026 · 2 citations
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan et al.CVPR 2022 · 66 citations
