Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
Yongchao Xu, Jiawei Liu, Junfeng Wang, Sen Tao, Na Jiang, Zheng-Jun Zha
Abstract
Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features to localize human-object pairs. However, such one-stage paradigms tend to overfit seen interactions, limiting their generalization to unseen categories, while the coarse spatial awareness of CLIP also hinders the localization of fine-grained interaction cues. To address these issues, we propose a novel Semantic-Diversified and Interaction-Focused framework (SD-IF), which integrates reinforcement-guided adaptive optimization to jointly enhance semantic generalization and spatial discrimination. Specifically, we introduce a Semantic Diversification (SD) module that applies reinforcementdriven stochastic semantic perturbations and dual-level semantic exploration, expanding the semantic coverage of queries while maintaining visual coherence and effectively encouraging exploration beyond the seen semantic clusters. Furthermore, we design an Interaction Focusing (IF) module that formulates an actor-critic optimization scheme to adaptively refine attention distributions based on detection features and interaction representations, guided by a hybrid reward combining spatial focusing and semantic consistency. Extensive experiments on two standard benchmarks demonstrate the effectiveness of our SD-IF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea24810-5c2a-4e52-8d34-2e8a73ad3608Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan et al.ACM MM 2024 · 5 citations
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsShan Ning, Longtian Qiu, Yongfei Liu, Xuming HeCVPR 2023
- Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-RobustnessGuangzhi Wang, Yangyang Guo, Ziwei Xu, Mohan S. KankanhalliCVPR 2024 · 9 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie et al.AAAI 2026 · 3 citations
