Look Less Think More: Rethinking Compositional Action Recognition
Rui Yan, Peng Huang, Xiangbo Shu, Junhao Zhang, Yonghua Pan, Jinhui Tang
摘要
Compositional action recognition which aims to identify the unseen combinations of actions and objects has recently attracted wide attention. Conventional methods bring in additional cues (e.g., dynamic motions of objects) to alleviate the inductive bias between the visual appearance of objects and the human action-level labels. Besides, compared with non-compositional settings, previous methods only pursue higher performance in compositional settings, which can not prove their generalization ability. To this end, we firstly rethink the problem and design a more generalized metric (namely Drop Ratio) and a more practical setting to evaluate the compositional generalization of existing action recognition algorithms. Beyond that, we propose a simple yet effective framework, Look Less Think More (LLTM), to reduce the strong association between visual objects and action-level labels (Look Less), and then discover the commonsense relationships between object categories and human actions (Think More). We test the rationality of the proposed Drop Ratio and Practical setting by comparing several popular action recognition methods on SSV2. Besides, the proposed LLTM achieves state-of-the-art performance on SSV2 with different settings.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Chop & Learn: Recognizing and Generating Object-State CompositionsNirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara 等ICCV 2023 · 被引用 20 次
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu 等CVPR 2024 · 被引用 12 次
- Improving Human-Object Interaction Detection via Virtual Image LearningShuman Fang, Shuai Liu, Jie Li, Guannan Jiang 等ACM MM 2023 · 被引用 6 次
- Zero-shot Compositional Action Recognition with Neural Logic ConstraintsGefan Ye, Lin Li, Kexin Li, Jun Xiao 等ACM MM 2025 · 被引用 1 次
相关 Paper
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu 等CVPR 2020
- Counterfactual Debiasing Inference for Compositional Action RecognitionPengzhan Sun, Bo Wu, Xunsong Li, Wen Li 等ACM MM 2021 · 被引用 25 次
- Motion Guided Attention Fusion to Recognize Interactions from VideosTae Soo Kim, Jonathan D. Jones, Gregory D. HagerICCV 2021 · 被引用 19 次
- VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict EntailmentDarshana Saravanan, Varun Gupta, Darshan Singh S, Zeeshan Khan 等CVPR 2025
- Watch Less, Do More: Implicit Skill Discovery for Video-Conditioned PolicyJiangxing Wang, Zongqing LuICLR 2025
