Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint Loss
Yuxiao Wang, Yu Lei, Zhenao Wei, Weiying Xue, Xinyu Jiang, Nan Zhuang, Qi Liu
Abstract
The task of Human-Object conTact (HOT) detection involves identifying the specific areas of the human body that are touching objects. Nevertheless, current models are restricted to just one type of image, often leading to too much segmentation in areas with little interaction, and struggling to maintain category consistency within specific regions. To tackle this issue, a HOT framework, termed P3HOT, is proposed, which blends Prompt guidance and human Proximal Perception. To begin with, we utilize a semantic-driven prompt mechanism to direct the network's attention towards the relevant regions based on the correlation between image and text. Then a human proximal perception mechanism is employed to dynamically perceive key depth range around the human, using learnable parameters to effectively eliminate regions where interactions are not expected. Calculating depth resolves the uncertainty of the overlap between humans and objects in a 2D perspective, providing a quasi-3D viewpoint. Moreover, a Regional Joint Loss (RJLoss) has been created as a new loss to inhibit abnormal categories in the same area. A new evaluation metric called ``AD-Acc.'' is introduced to address the shortcomings of existing methods in addressing negative samples. Comprehensive experimental results demonstrate that our approach achieves state-of-the-art performance in four metrics across two benchmark datasets. Specifically, our model achieves an improvement of 0.7, 2.0, 1.6, and 11.0 in SC-Acc., mIoU, wIoU, and AD-Acc. metrics, respectively, on the HOT-Annotated dataset. The sources code are available at https://github.com/YuxiaoWang-AI/P3HOT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5d46c10-0a9b-48dc-b78a-74d579d7f200Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- BEHAVE: Dataset and Method for Tracking Human Object InteractionsBharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu et al.CVPR 2022 · 144 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Detecting Hands and Recognizing Physical Contact in the WildSupreeth Narasimhaswamy, Trung Nguyen, Minh Hoai NguyenNeurIPS 2020 · 57 citations
Related papers
- Precision-Enhanced Human-Object Contact Detection via Depth-Aware Perspective Interaction and Object Texture RestorationYuxiao Wang, Wenpeng Neng, Zhenao Wei, Yu Lei et al.AAAI 2025 · 6 citations
- Detecting Human-Object Contact in ImagesYixin Chen, Sai Kumar Dwivedi, Michael J. Black, Dimitrios TzionasCVPR 2023
- HORP: Human-Object Relation Priors Guided HOI DetectionPei Geng, Jian Yang, Shanshan ZhangCVPR 2025
- HOTR: End-to-End Human-Object Interaction Detection With TransformersBumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim et al.CVPR 2021
- Deep Contextual Attention for Human-Object Interaction DetectionTiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan et al.ICCV 2019 · 130 citations
