What-Meets-Where: Unified Learning of Action and Contact Localization in Images
Yuxiao Wang, Yu Lei, Wolin Liang, Weiying Xue, Zhenao Wei, Nan Zhuang, Qi Liu
Abstract
People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider what action is occurring and where it is happening. Current methodologies, however, often inadequately capture this duality, typically failing to jointly model both action semantics and their spatial contextualization within scenes. To bridge this gap, we introduce a novel vision task that simultaneously predicts high-level action semantics and fine-grained body-part contact regions. Our proposed framework, PaIR-Net, comprises three key components: the Contact Prior Aware Module (CPAM) for identifying contact-relevant body parts, the Prior-Guided Concat Segmenter (PGCS) for pixel-wise contact segmentation, and the Interaction Inference Module (IIM) responsible for integrating global interaction relationships. To facilitate this task, we present PaIR (Part-aware Interaction Representation), a comprehensive dataset containing 13,979 images that encompass 654 actions, 80 object categories, and 17 body parts. Experimental evaluation demonstrates that PaIR-Net significantly outperforms baseline approaches, while ablation studies confirm the efficacy of each architectural component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0016c4b-cbdb-4a44-9ae9-0f33d7b43f78Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Capturing and Inferring Dense Full-Body Human-Scene ContactChun-Hao P. Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin et al.CVPR 2022 · 106 citations
- Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction DetectionYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2022 · 88 citations
- Assessing Depth Perception in VR and Video See-Through AR: A Comparison on Distance Judgment, Performance, and PreferenceFranziska Westermeier, Larissa Brübach, Carolin Wienrich, Marc Erich LatoschikIEEE VR 2024 · 38 citations
Related papers
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li et al.ICCV 2019 · 224 citations
- DECO: Dense Estimation of 3D Human-Scene Contact In The WildShashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi et al.ICCV 2023 · 54 citations
- HOPNet: Learning Hand-Object-Person Interaction Network for Hand Contact State DetectionWei Li, Yizhao Wan, Xiao Wu, Jianshuai Wang et al.ACM MM 2025 · 1 citation
- Detecting Human-Object Contact in ImagesYixin Chen, Sai Kumar Dwivedi, Michael J. Black, Dimitrios TzionasCVPR 2023
- Leverage Interactive Affinity for Affordance LearningHongchen Luo, Wei Zhai, Jing Zhang, Yang Cao et al.CVPR 2023
