HORP: Human-Object Relation Priors Guided HOI Detection
Pei Geng, Jian Yang, Shanshan Zhang
Abstract
Human-Object Interaction (HOI) detection aims to predict the <Human, Interaction, Object> triplets, where the core challenge lies in recognizing the interaction of each humanobject pair. Despite recent progress thanks to more advanced model architectures, HOI performance remains unsatisfactory. In this work, we first perform some failure analysis and find that the accuracy of the no-interaction category is extremely low, largely hindering the improvement of overall performance. We further look into the error types and find the mis-classification between no-interaction and with-interaction ones can be handled by human-object relation priors. Specifically, to better distinguish no-interaction from direct interactions, we propose 3D location prior, which indicates the distance between human and object; as of no-interaction vs. indirect interactions, we propose gaze area prior, which denotes whether human can see the object or not. The above two types of human-object relation priors are represented by text and are combined with the original visual features, generating multi-modal cues for interaction recognition. Experimental results on the HICO-DET and V-COCO datasets demonstrate that our proposed human-object relation priors are effective and our method HORP surpasses previous methods under various settings and scenarios. In particular, the usage of our priors significantly enhances the model's recognition ability for the no-interaction category. Code is available at https://github.com/namegp/HORP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea567209-b9b4-4959-8714-5245754d2ef6Cited by top-tier papers2
- Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction DetectionSoo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son et al.CVPR 2026 · 1 citation
- LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction DetectionEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangICLR 2026
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu et al.NeurIPS 2021 · 218 citations
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 155 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
Related papers
- Learning Human-Object Interaction Detection Using Interaction PointsTiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan et al.CVPR 2020
- Distance Matters in Human-Object Interaction DetectionGuangzhi Wang, Yangyang Guo, Yongkang Wong, Mohan S. KankanhalliACM MM 2022 · 18 citations
- Deep Contextual Attention for Human-Object Interaction DetectionTiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan et al.ICCV 2019 · 130 citations
- Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction DetectionWeibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu et al.AAAI 2024 · 11 citations
- End-to-End Human Object Interaction Detection With HOI TransformerCheng Zou, Bohan Wang, Yue Hu, Junqi Liu et al.CVPR 2021
