Agglomerative Transformer for Human-Object Interaction Detection
Danyang Tu, Wei Sun, Guangtao Zhai, Wei Shen
Abstract
We propose an agglomerative Transformer (AGER) that enables Transformer-based human-object interaction (HOI) detectors to flexibly exploit extra instance-level cues in a single-stage and end-to-end manner for the first time. AGER acquires instance tokens by dynamically clustering patch tokens and aligning cluster centers to instances with textual guidance, thus enjoying two benefits: 1) Integrality: each instance token is encouraged to contain all discriminative feature regions of an instance, which demonstrates a significant improvement in the extraction of different instance-level cues and subsequently leads to a new state-of-the-art performance of HOI detection with 36.75 mAP on HICO-Det. 2) Efficiency: the dynamical clustering mechanism allows AGER to generate instance tokens jointly with the feature learning of the Transformer encoder, eliminating the need of an additional object detector or instance decoder in prior methods, thus allowing the extraction of desirable extra cues for HOI detection in a single-stage and end-to-end pipeline. Concretely, AGER reduces GFLOPs by 8.5% and improves FPS by 36%, even compared to a vanilla DETR-like pipeline without extra cue extraction. The code will be available at https://github.com/six6607/AGER.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 42 citations
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionYuchen Zhou, Linkai Liu, Chao GouCVPR 2024 · 13 citations
- Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen et al.CVPR 2024 · 10 citations
- Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-RobustnessGuangzhi Wang, Yangyang Guo, Ziwei Xu, Mohan S. KankanhalliCVPR 2024 · 9 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Anchor DETR: Query Design for Transformer-Based DetectorYingming Wang, Xiangyu Zhang, Tong Yang, Jian SunAAAI 2022 · 567 citations
Related papers
- End-to-End Human Object Interaction Detection With HOI TransformerCheng Zou, Bohan Wang, Yue Hu, Junqi Liu et al.CVPR 2021
- QPIC: Query-Based Pairwise Human-Object Interaction Detection With Image-Wide Contextual InformationMasato Tamura, Hiroki Ohashi, Tomoaki YoshinagaCVPR 2021
- Reformulating HOI Detection As Adaptive Set PredictionMingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen et al.CVPR 2021
- Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction DetectionYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2022 · 88 citations
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
