Interaction Compass: Multi-Label Zero-Shot Learning of Human-Object Interactions via Spatial Relations
Dat Huynh, Ehsan Elhamifar
摘要
We study the problem of multi-label zero-shot recognition in which labels are in the form of human-object interactions (combinations of actions on objects), each image may contain multiple interactions and some interactions do not have training images. We propose a novel compositional learning framework that decouples interaction labels into separate action and object scores that incorporate the spatial compatibility between the two components. We combine these scores to efficiently recognize seen and unseen interactions. However, learning action-object spatial relations, in principle, requires bounding-box annotations, which are costly to gather. Moreover, it is not clear how to generalize spatial relations to unseen interactions. We address these challenges by developing a cross-attention mechanism that localizes objects from action locations and vice versa by predicting displacements between them, referred to as relational directions. During training, we estimate the relational directions as ones maximizing the scores of ground-truth interactions that guide predictions toward compatible action-object regions. By extensive experiments, we show the effectiveness of our framework, where we improve the state of the art by 2.6% mAP score and 5.8% recall score on HICO and Visual Genome datasets, respectively.1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingDat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu 等CVPR 2022 · 被引用 78 次
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan 等CVPR 2022 · 被引用 66 次
- An Image-like Diffusion Method for Human-Object Interaction DetectionXiaofei Hui, Haoxuan Qu, Hossein Rahmani, Jun LiuCVPR 2025
- Compositional Targeted Multi-Label Universal PerturbationsHassan Mahmood, Ehsan ElhamifarCVPR 2025
它引用的顶会 Paper23
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu 等ICCV 2019 · 被引用 347 次
- Transferable Contrastive Network for Generalized Zero-Shot LearningHuajie Jiang, Ruiping Wang, Shiguang Shan, Xilin ChenICCV 2019 · 被引用 200 次
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 被引用 155 次
- No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training TechniquesTanmay Gupta, Alexander G. Schwing, Derek HoiemICCV 2019 · 被引用 149 次
- Detecting Unseen Visual Relations Using AnalogiesJulia Peyre, Josef Sivic, Ivan Laptev, Cordelia SchmidICCV 2019 · 被引用 135 次
相关 Paper
- A Shared Multi-Attention Framework for Multi-Label Zero-Shot LearningDat Huynh, Ehsan ElhamifarCVPR 2020
- Discovering Human Interactions With Novel Objects via Zero-Shot LearningSuchen Wang, Kim-Hui Yap, Junsong Yuan, Yap-Peng TanCVPR 2020
- End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationMingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin 等AAAI 2023 · 被引用 64 次
- Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction RecognitionShiyu Xuan, Dongkai Wang, Zechao Li, Jinhui TangICLR 2026 · 被引用 2 次
- ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction DetectionYe Liu, Junsong Yuan, Chang Wen ChenACM MM 2020 · 被引用 83 次
