Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction Networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, Trevor Darrell
摘要
Human action is naturally compositional: humans can easily recognize and perform actions with objects that are different from those used in training demonstrations. In this paper, we study the compositionality of action by looking into the dynamics of subject-object interactions. We propose a novel model which can explicitly reason about the geometric relations between constituent objects and an agent performing an action. To train our model, we collect dense object box annotations on the Something-Something dataset. We propose a novel compositional action recognition task where the training combinations of verbs and nouns do not overlap with the test set. The novel aspects of our model are applicable to activities with prominent object interaction dynamics and to objects which can be tracked using state-of-the-art approaches; for activities without clearly defined spatial object-agent interactions, we rely on baseline scene-level spatio-temporal representations. We show the effectiveness of our approach not only on the proposed compositional action recognition task, but also in a few-shot compositional setting which requires the model to generalize across both object appearance and action category. 1 * Equal advising 1 Project page: https://joaanna.github.io/something else/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig 等NeurIPS 2023 · 被引用 93 次
- ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and SynthesisLixin Yang, Kailin Li, Xinyu Zhan, Jun Lv 等CVPR 2022 · 被引用 82 次
- Object-Region Video TransformersRoei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar 等CVPR 2022 · 被引用 75 次
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu 等CVPR 2022 · 被引用 63 次
- Implicit Temporal Modeling with Learnable Alignment for Video RecognitionShuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng 等ICCV 2023 · 被引用 63 次
它引用的顶会 Paper3
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Reasoning About Human-Object Interactions Through Dual Attention NetworksTete Xiao, Quanfu Fan, Danny Gutfreund, Mathew Monfort 等ICCV 2019 · 被引用 36 次
- Few-Shot Video Classification via Temporal AlignmentKaidi Cao, Jingwei Ji, Zhangjie Cao, Chien-Yi Chang 等CVPR 2020
相关 Paper
- Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using SuperquadricsTze Ho Elden Tse, Runyang Feng, Linfang Zheng, Jiho Park 等AAAI 2025 · 被引用 2 次
- Counterfactual Debiasing Inference for Compositional Action RecognitionPengzhan Sun, Bo Wu, Xunsong Li, Wen Li 等ACM MM 2021 · 被引用 25 次
- Motion Guided Attention Fusion to Recognize Interactions from VideosTae Soo Kim, Jonathan D. Jones, Gregory D. HagerICCV 2021 · 被引用 19 次
- Look Less Think More: Rethinking Compositional Action RecognitionRui Yan, Peng Huang, Xiangbo Shu, Junhao Zhang 等ACM MM 2022 · 被引用 17 次
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 被引用 28 次
