LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection
Eastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin Wang
Abstract
Human-Object Interaction (HOI) detection with vision-language models (VLMs) has progressed rapidly, yet a trade-off persists between specialization and generalization. Two major challenges remain: (1) the sparsity of supervision, which hampers effective transfer of foundation models to HOI tasks, and (2) the absence of a generalizable architecture that can excel in both fully supervised and zero-shot scenarios. To address these issues, we propose LINK, Learning INstance-level Knowledge from VLMs. First, we introduce a HOI detection framework equipped with a Human-Object Geometrical Encoder and a VLM Linking Decoder. By decoupling from detector-specific features, our design ensures plug-and-play compatibility with arbitrary object detectors and consistent adaptability across diverse settings. Building on this foundation, we develop a Progressive Learning Strategy under a teacher-student paradigm, which delivers dense supervision over all potential human-object pairs. By contrasting subtle spatial and semantic differences between positive and negative instances, the model learns robust and transferable HOI representations. LINK sets new state-of-the-art on SWiG-HOI, HICO-DET, and V-COCO across zero-shot, fully supervised, and open-vocabulary settings, with strong generalization to unseen interaction categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbbf4cdd-386b-4f60-b3b6-d733f79f0b65Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- GLIPv2: Unifying Localization and Vision-Language UnderstandingHaotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen et al.NeurIPS 2022 · 403 citations
Related papers
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationQinqian Lei, Bo Wang, Robby T. TanICCV 2025 · 4 citations
- End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationMingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin et al.AAAI 2023 · 64 citations
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 42 citations
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li et al.NeurIPS 2023 · 62 citations
- Efficient Adaptive Human-Object Interaction Detection with Concept-guided MemoryTing Lei, Fabian Caba, Qingchao Chen, Hailin Jin et al.ICCV 2023 · 57 citations
