RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
Jihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song, Hyunwoo J. Kim
Abstract
Weakly supervised Human–Object Interaction (HOI) detection is vital for scalable scene understanding by learning interactions from only image-level annotations, i.e., no labels specifying which human–object instances are engaged in the interaction.Due to the lack of localization signals, prior works typically propose candidate pairs using an external object detector and then infer their interactions through pairwise reasoning.However, this framework often struggles to scale due to the substantial computational cost incurred by enumerating numerous instance pairs. In addition, it exhibits suboptimal performance due to false positives arising from non-interactive combinations, hindering its capability of instance-level HOI reasoning.To this end, we introduce Relational Grounding Transformer (RegFormer), a versatile interaction recognition module that enables efficient and accurate HOI reasoning.Under image-level supervision, RegFormer leverages spatially grounded implicit signals as guidance for the reasoning process, facilitating effective locality elicitation.Benefiting from the implicitly learned local interactions, our module can accurately distinguish humans, objects, and their interactions within their corresponding regions, enabling precise and efficient instance-level HOI reasoning without any additional training.Our extensive experiments and analysis demonstrate that RegFormer effectively learns spatial cues for instance-level interaction reasoning, operates with high efficiency, and even shows comparable performance compared to fully supervised models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b62b0593-0e7b-464c-9c44-83a62105aae7Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerFrederic Z. Zhang, Dylan Campbell, Stephen GouldCVPR 2022 · 118 citations
- RLIP: Relational Language-Image Pre-training for Human-Object Interaction DetectionHangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng et al.NeurIPS 2022 · 88 citations
- Exploring Predicate Visual Context in Detecting of Human-Object InteractionsFrederic Z. Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong et al.ICCV 2023 · 86 citations
Related papers
- HOTR: End-to-End Human-Object Interaction Detection With TransformersBumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim et al.CVPR 2021
- End-to-End Human Object Interaction Detection With HOI TransformerCheng Zou, Bohan Wang, Yue Hu, Junqi Liu et al.CVPR 2021
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
- MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction DetectionBumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin et al.CVPR 2022 · 80 citations
- QPIC: Query-Based Pairwise Human-Object Interaction Detection With Image-Wide Contextual InformationMasato Tamura, Hiroki Ohashi, Tomoaki YoshinagaCVPR 2021
