MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction Detection
Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin, Junhyun Lee, Eun-Sol Kim
Abstract
Human-Object Interaction (HOI) detection is the task of identifying a set of (human, object, interaction) triplets from an image. Recent work proposed transformer encoder-decoder architectures that successfully eliminated the need for many hand-designed components in HOI detection through end-to-end training. However, they are limited to single-scale feature resolution, providing suboptimal performance in scenes containing humans, objects, and their interactions with vastly different scales and distances. To tackle this problem, we propose a Multi-Scale TRansformer (MSTR) for HOI detection powered by two novel HOI-aware deformable attention modules called Dual-Entity attention and Entity-conditioned Context attention. While existing deformable attention comes at a huge cost in HOI detection performance, our proposed attention modules of MSTR learn to effectively attend to sampling points that are essential to identify interactions. In experiments, we achieve the new state-of-the-art performance on two HOI detection benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d516f4a4-3419-41c0-804b-a9cef541d80fCited by top-tier papers29
- Exploring Predicate Visual Context in Detecting of Human-Object InteractionsFrederic Z. Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong et al.ICCV 2023 · 86 citations
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie et al.ICCV 2023 · 69 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- Neural-Logic Human-Object Interaction DetectionLiulei Li, Jianan Wei, Wenguan Wang, Yi YangNeurIPS 2023 · 54 citations
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li et al.ICCV 2019 · 224 citations
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 155 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
- No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training TechniquesTanmay Gupta, Alexander G. Schwing, Derek HoiemICCV 2019 · 149 citations
Related papers
- HOTR: End-to-End Human-Object Interaction Detection With TransformersBumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim et al.CVPR 2021
- Human-Object Interaction Detection via Disentangled TransformerDesen Zhou, Zhichao Liu, Jian Wang, Leshan Wang et al.CVPR 2022 · 62 citations
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
- Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction DetectionYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2022 · 88 citations
- QPIC: Query-Based Pairwise Human-Object Interaction Detection With Image-Wide Contextual InformationMasato Tamura, Hiroki Ohashi, Tomoaki YoshinagaCVPR 2021
