Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
Soo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son, Nam Ik Cho, Jun Won Choi
摘要
Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have leveraged Vision-Language Models (VLMs) to introduce semantic priors, significantly improving HOI detection performance. However, existing methods often fail to fully capitalize on the diverse contextual cues distributed across the entire scene. To overcome these limitations, we propose the Instance-centric Context Mining Network (InCoM-Net)-a novel framework that effectively integrates rich semantic knowledge extracted from VLMs with instance-specific features produced by an object detector. This design enables deeper interaction reasoning by modeling relationships not only within each detected instance but also across instances and their surrounding scene context. InCoM-Net comprises two core components: Instancecentric Context Refinement (ICR), which separately extracts intra-instance, inter-instance, and global contextual cues from VLM-derived features, and Progressive Context Aggregation (ProCA), which iteratively fuses these multicontext features with instance-level detector features to support high-level HOI reasoning. Extensive experiments on the HICO-DET and V-COCO benchmarks show that InCoM-Net achieves state-of-the-art performance, surpassing previous HOI detection methods. Code is available at https://github.com/nowuss/InCoM-Net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Spatially Conditioned Graphs for Detecting Human-Object InteractionsFrederic Z. Zhang, Dylan Campbell, Stephen GouldICCV 2021 · 被引用 170 次
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang 等CVPR 2022 · 被引用 136 次
相关 Paper
- Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen 等CVPR 2024 · 被引用 10 次
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionYupeng Hu, Changxing Ding, Chang Sun, Shaoli Huang 等ICCV 2025 · 被引用 1 次
- Learning Human-Object Interaction Detection Using Interaction PointsTiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan 等CVPR 2020
- Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction DetectionFrancesco Tonini, Lorenzo Vaquero, Alessandro Conti, Cigdem Beyan 等ACM MM 2025 · 被引用 2 次
- LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction DetectionEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangICLR 2026
