Learning Transferable Human-Object Interaction Detector with Natural Language Supervision
Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan
摘要
It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey 等CHI 2024 · 被引用 124 次
- Teachers, Parents, and Students' perspectives on Integrating Generative AI into Elementary Literacy EducationAriel Han, Xiaofei Zhou, Zhenyao Cai, Shenshen Han 等CHI 2024 · 被引用 104 次
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen 等NeurIPS 2023 · 被引用 64 次
- End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationMingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin 等AAAI 2023 · 被引用 64 次
- CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI DetectionYunyao Mao, Jiajun Deng, Wengang Zhou, Li Li 等NeurIPS 2023 · 被引用 62 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li 等ICCV 2019 · 被引用 224 次
相关 Paper
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsShan Ning, Longtian Qiu, Yongfei Liu, Xuming HeCVPR 2023
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang 等CVPR 2022 · 被引用 136 次
- Discovering Human Interactions with Large-Vocabulary Objects via Query and Multi-Scale DetectionSuchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu 等ICCV 2021 · 被引用 35 次
- LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction DetectionEastman Z. Y. Wu, Yali Li, Yuan Wang, Shengjin WangICLR 2026
- Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI DetectionYixin Guo, Yu Liu, Jianghao Li, Weimin Wang 等ACM MM 2024 · 被引用 12 次
