Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection
Jiayi Gao, Kongming Liang, Tao Wei, Wei Chen, Zhanyu Ma, Jun Guo
Abstract
Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f02eda4f-66ca-4154-b345-bd7d79fe27a8Cited by top-tier papers3
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen et al.ICCV 2025 · 5 citations
- HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI DetectionYongchao Xu, Jiawei Liu, Sen Tao, Qiang Zhang et al.AAAI 2025 · 1 citation
- LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction DetectionEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangICLR 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
Related papers
- Improving Human-Object Interaction Detection via Phrase Learning and Label CompositionZhimin Li, Cheng Zou, Yu Zhao, Boxun Li et al.AAAI 2022 · 43 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Disentangled Pre-Training for Human-Object Interaction DetectionZhuolong Li, Xingao Li, Changxing Ding, Xiangmin XuCVPR 2024
- ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction DetectionYe Liu, Junsong Yuan, Chang Wen ChenACM MM 2020 · 83 citations
- Detecting Human-Object Interaction via Fabricated Compositional LearningZhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng et al.CVPR 2021
