Open-World Human-Object Interaction Detection via Multi-Modal Prompts
Jie Yang, Bingliang Li, Ailing Zeng, Lei Zhang, Ruimao Zhang
摘要
In this paper, we develop MP-HOI, a powerful Multimodal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions, realizing HOI detection in the open world. Specifically, it integrates visual prompts into existing language-guided-only HOI detectors to handle situations where textual descriptions face difficulties in generalization and to address complex scenarios with high interaction ambiguity. To facilitate MP-HOI training, we build a large-scale HOI dataset named Magic-HOI, which gathers six existing datasets into a unified label space, forming over 186K images with 2.4K objects, 1.2K actions, and 20K HOI interactions. Furthermore, to tackle the long-tail issue within the Magic-HOI dataset, we introduce an automated pipeline for generating realistically annotated HOI images and present SynHOI, a high-quality synthetic HOI dataset containing 100K images. Leveraging these two datasets, MP-HOI optimizes the HOI task as a similarity learning process between multi-modal prompts and objects/interactions via a unified contrastive loss, to learn generalizable and transferable objects/interactions representations from large-scale data. MP-HOI could serve as a generalist HOI detector, surpassing the HOI vocabulary of existing expert models by more than 30 times. Concurrently, our results demonstrate that MP-HOI exhibits remarkable zero-shot capability in real- world scenarios and consistently achieves a new state-of-the-art performance across various benchmarks. Our project homepage is available at https://MP-HOI.github.iol.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng 等ICCV 2025 · 被引用 6 次
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun 等ICCV 2025 · 被引用 6 次
- CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific MethodsQinqian Lei, Bo Wang, Robby T. TanCVPR 2026 · 被引用 6 次
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI DetectionChanhyeong Yang, Taehoon Song, Jihwan Park, Hyunwoo J. KimNeurIPS 2025 · 被引用 5 次
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationQinqian Lei, Bo Wang, Robby T. TanICCV 2025 · 被引用 4 次
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- Toward Open-Set Human Object Interaction DetectionMingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun 等AAAI 2024 · 被引用 12 次
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang 等NeurIPS 2025 · 被引用 3 次
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan 等CVPR 2022 · 被引用 66 次
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 被引用 29 次
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen 等NeurIPS 2023 · 被引用 64 次
