Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
Ting Lei, Shaofeng Yin, Yang Liu
Abstract
Open-vocabulary human-object interaction (HOI) detection, which is concerned with the problem of detecting novel HOIs guided by natural language, is crucial for understanding human-centric scenes. However, prior zeroshot HOI detectors often employ the same levels of feature maps to model HOIs with varying distances, leading to suboptimal performance in scenes containing humanobject pairs with a wide range of distances. In addition, these detectors primarily rely on category names and overlook the rich contextual information that language can provide, which is essential for capturing open vocabulary concepts that are typically rare and not well-represented by category names alone. In this paper, we introduce a novel end-to-end open vocabulary HOI detection framework with conditional multi-level decoding and fine-grained semantic enhancement (CMD-SE), harnessing the potential of Visual-Language Models (VLMs). Specifically, we propose to model human-object pairs with different distances with different levels of feature maps by incorporating a soft constraint during the bipartite matching process. Furthermore, by leveraging large language models (LLMs) such as GPT models, we exploit their extensive world knowledge to generate descriptions of human body part states for various interactions. Then we integrate the generalizable and fine-grained semantics of human body parts to improve interaction recognition. Experimental results on two datasets, SWIG-HOI and HICO-DET, demonstrate that our proposed method achieves state-of-the-art results in open vocabulary HOI detection. The code and models are available at https://github.com/ltttpku/CMD-SErelease .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0ffaa78-9587-4e89-9e35-f2f6914c49b9Cited by top-tier papers21
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 42 citations
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- PlanLLM: Video Procedure Planning with Refinable Large Language ModelsDejie Yang, Zijing Zhao, Yang LiuAAAI 2025 · 8 citations
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng et al.ICCV 2025 · 6 citations
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingYUEJIAO SU, Yi Wang, Lei Yao, Yawen Cui et al.ICLR 2026 · 5 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- What does a platypus look like? Generating customized prompts for zero-shot image classificationSarah M. Pratt, Ian Covert, Rosanne Liu, Ali FarhadiICCV 2023 · 343 citations
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu et al.NeurIPS 2021 · 218 citations
- Spatially Conditioned Graphs for Detecting Human-Object InteractionsFrederic Z. Zhang, Dylan Campbell, Stephen GouldICCV 2021 · 170 citations
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 155 citations
Related papers
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction RecognitionShiyu Xuan, Dongkai Wang, Zechao Li, Jinhui TangICLR 2026 · 2 citations
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionYupeng Hu, Changxing Ding, Chang Sun, Shaoli Huang et al.ICCV 2025 · 1 citation
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan et al.ACM MM 2024 · 5 citations
- Streamlined Open-Vocabulary Human-Object Interaction DetectionChang Sun, Dongliang Liao, Changxing DingCVPR 2026 · 2 citations
