Reasoning About Human-Object Interactions Through Dual Attention Networks
Tete Xiao, Quanfu Fan, Danny Gutfreund, Mathew Monfort, Aude Oliva, Bolei Zhou
摘要
Objects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reasons about human-object interactions. The dualattentional framework weights the important features for objects and actions respectively. As a result, the recognition of objects and actions mutually benefit each other. The proposed model shows competitive classification performance on the human-object interaction dataset Something-Something. Besides, it can perform weak spatiotemporal localization and affordance segmentation, despite being trained only with video-level labels. The model not only finds when an action is happening and which object is being manipulated, but also identifies which part of the object is being interacted with. Project page: https: //dual-attention-network.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals MeasurementXin Liu, Josh Fromm, Shwetak N. Patel, Daniel McDuffNeurIPS 2020 · 被引用 436 次
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo 等ICCV 2021 · 被引用 271 次
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan 等CVPR 2022 · 被引用 66 次
- Human-Object Interaction Detection via Disentangled TransformerDesen Zhou, Zhichao Liu, Jian Wang, Leshan Wang 等CVPR 2022 · 被引用 62 次
- Discovering Human Interactions with Large-Vocabulary Objects via Query and Multi-Scale DetectionSuchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu 等ICCV 2021 · 被引用 35 次
相关 Paper
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu 等CVPR 2020
- Motion Guided Attention Fusion to Recognize Interactions from VideosTae Soo Kim, Jonathan D. Jones, Gregory D. HagerICCV 2021 · 被引用 19 次
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction RecognitionNing Wang, Guangming Zhu, Liang Zhang, Peiyi Shen 等ACM MM 2021 · 被引用 32 次
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
