Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware Queries
Zhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan, Dongyang Zhang, Tao He, Yuan-Fang Li
Abstract
The open-vocabulary human-object interaction (Ov-HOI) detection aims to identify both base and novel categories of human-object interactions while only base categories are available during training. Existing Ov-HOI methods commonly leverage knowledge distilled from CLIP to extend their ability to detect previously unseen interaction categories. However, our empirical observations indicate that the inherent noise present in CLIP has a detrimental effect on HOI prediction. Moreover, the absence of novel human-object position distributions often leads to overfitting on the base categories within their learned queries. To address these issues, we propose a two-step framework named, CaM-LQ, Calibrating visual-language Models, (e.g., CLIP) for open-vocabulary HOI detection with Locality-aware Queries. By injecting the fine-grained HOI supervision from the calibrated CLIP into the HOI decoder, our model can achieve the goal of predicting novel interactions. Extensive experimental results demonstrate that our approach performs well in open-vocabulary human-object interaction detection, surpassing state-of-the-art methods across multiple metrics on mainstream datasets and showing superior open-vocabulary HOI detection performance, e.g., with 4.54 points improvement on the HICO-DET dataset over the SoTA CLIP4HOI on the UV task with the same backbone ResNet-50.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7cc62e3f-0702-4556-b819-9142cbdd2489Cited by top-tier papers2
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng et al.ICCV 2025 · 6 citations
- TiCAL: Typicality-Based Consistency-Aware Learning for Multimodal Emotion RecognitionWen Yin, Siyu Zhan, Cencen Liu, Xin Hu et al.AAAI 2026 · 4 citations
Related papers
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsShan Ning, Longtian Qiu, Yongfei Liu, Xuming HeCVPR 2023
- Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI DetectionYongchao Xu, Jiawei Liu, Junfeng Wang, Sen Tao et al.CVPR 2026
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan et al.CVPR 2022 · 66 citations
- Weakly-supervised HOI Detection via Prior-guided Bi-level Representation LearningBo Wan, Yongfei Liu, Desen Zhou, Tinne Tuytelaars et al.ICLR 2023 · 5 citations
