Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
Yu-Wei Zhan, Fan Liu, Xin Luo, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli
Abstract
Human-Object Interaction (HOI) detection aims at detecting human-object pairs and predicting their interactions. However, conventional HOI detection methods often struggle to fully capture the contextual information needed to accurately identify these interactions. While large Vision-Language Models (VLMs) show promise in tasks involving human interactions, they are not tailored for HOI detection. The complexity of human behavior and the diverse contexts in which these interactions occur make it further challenging. Contextual cues, such as the participants involved, body language, and the surrounding environment, play crucial roles in predicting these interactions, especially those that are unseen or ambiguous. Moreover, large VLMs are trained on vast image and text data, enabling them to generate contextual cues that help in understanding real-world contexts, object relationships, and typical interactions. Building on this, in this paper we introduce ConCue, a novel approach for improving visual feature extraction in HOI detection. Specifically, we first design specialized prompts to utilize large VLMs to generate contextual cues within an image. To fully leverage these cues, we develop a transformer-based feature extraction module with a multi-tower architecture that integrates contextual cues into both instance and interaction detectors. Extensive experiments and analyses demonstrate the effectiveness of using these contextual cues for HOI detection. The experimental results show that integrating ConCue with existing state-of-the-art methods significantly enhances their performance on two widely used datasets. The code is available at https://github.com/yw- zhan/ConCue.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fc0bfbc-9f70-4e33-894c-e0bf991d17acCited by top-tier papers2
- Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationHaoyu Tang, Tianyuan Liang, Han Jiang, Xuesong Liu et al.AAAI 2026
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng et al.CVPR 2026
Builds on46
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li et al.ICCV 2019 · 224 citations
Related papers
- Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen et al.CVPR 2024 · 10 citations
- Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction DetectionSoo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son et al.CVPR 2026 · 1 citation
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI DetectionTing Lei, Shaofeng Yin, Yang LiuCVPR 2024
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionYupeng Hu, Changxing Ding, Chang Sun, Shaoli Huang et al.ICCV 2025 · 1 citation
