Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection
Jiayi Gao, Kongming Liang, Tao Wei, Wei Chen, Zhanyu Ma, Jun Guo
摘要
Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen 等ICCV 2025 · 被引用 5 次
- HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI DetectionYongchao Xu, Jiawei Liu, Sen Tao, Qiang Zhang 等AAAI 2025 · 被引用 1 次
- LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction DetectionEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangICLR 2026
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
相关 Paper
- Improving Human-Object Interaction Detection via Phrase Learning and Label CompositionZhimin Li, Cheng Zou, Yu Zhao, Boxun Li 等AAAI 2022 · 被引用 43 次
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang 等CVPR 2022 · 被引用 136 次
- Disentangled Pre-Training for Human-Object Interaction DetectionZhuolong Li, Xingao Li, Changxing Ding, Xiangmin XuCVPR 2024
- ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction DetectionYe Liu, Junsong Yuan, Chang Wen ChenACM MM 2020 · 被引用 83 次
- Detecting Human-Object Interaction via Fabricated Compositional LearningZhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng 等CVPR 2021
