Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise Transformer
Frederic Z. Zhang, Dylan Campbell, Stephen Gould
Abstract
Recent developments in transformer models for visual data have led to significant improvements in recognition and detection tasks. In particular, using learnable queries in place of region proposals has given rise to a new class of one-stage detection models, spearheaded by the Detection Transformer (DETR). Variations on this one-stage approach have since dominated human-object interaction (HOI) detection. However, the success of such one-stage HOI detectors can largely be attributed to the representation power of transformers. We discovered that when equipped with the same transformer, their two-stage counterparts can be more performant and memory-efficient, while taking a fraction of the time to train. In this work, we propose the Unary-Pairwise Transformer, a two-stage detector that exploits unary and pairwise representations for HOIs. We observe that the unary and pairwise parts of our transformer network specialise, with the former preferentially increasing the scores of positive examples and the latter decreasing the scores of negative examples. We evaluate our method on the HICO-DET and V-COCO datasets, and significantly outperform state-of-the-art approaches. At inference time, our model with ResNet50 approaches real-time performance on a single GPU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb84f627-88e3-47c6-9d1d-afc4f53c8bc0Cited by top-tier papers55
- Exploring Predicate Visual Context in Detecting of Human-Object InteractionsFrederic Z. Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong et al.ICCV 2023 · 86 citations
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie et al.ICCV 2023 · 69 citations
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- Efficient Adaptive Human-Object Interaction Detection with Concept-guided MemoryTing Lei, Fabian Caba, Qingchao Chen, Hailin Jin et al.ICCV 2023 · 57 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Spatially Conditioned Graphs for Detecting Human-Object InteractionsFrederic Z. Zhang, Dylan Campbell, Stephen GouldICCV 2021 · 170 citations
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen et al.EMNLP 2020 · 158 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
Related papers
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu et al.NeurIPS 2021 · 218 citations
- Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction DetectionYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2022 · 88 citations
- End-to-End Human Object Interaction Detection With HOI TransformerCheng Zou, Bohan Wang, Yue Hu, Junqi Liu et al.CVPR 2021
- QPIC: Query-Based Pairwise Human-Object Interaction Detection With Image-Wide Contextual InformationMasato Tamura, Hiroki Ohashi, Tomoaki YoshinagaCVPR 2021
