Further Understanding Videos through Adverbs: A New Video Task
Bo Pang, Kaiwen Zha, Yifan Zhang, Cewu Lu
摘要
Video understanding is a research hotspot of computer vision and significant progress has been made on video action recognition recently. However, the semantics information contained in actions is not rich enough to build powerful video understanding models. This paper first introduces a new video semantics: the Behavior Adverb (BA), which is a more expressive and difficult one covering subtle and inherent characteristics of human action behavior. To exhaustively decode this semantics, we construct the Videos with Action and Adverb Dataset (VAAD), which is a large-scale dataset with a semantically complete set of BAs. The dataset will be released to the public with this paper. We benchmark several representative video understanding methods (originally for action recognition) on BA and action recognition. The results show that BA recognition task is more challenging than conventional action recognition. Accordingly, we propose the BA Understanding Network (BAUN) to solve this problem and the experiments reveal that our BAUN is more suitable for BA recognition (11% better than I3D). Furthermore, we find these two semantics (action and BA) can propel each other forward to better performance: promoting action recognition results by 3.4% averagely on three standard action recognition datasets (UCF-101, HMDB-51, Kinetics).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li 等NeurIPS 2020 · 被引用 152 次
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Quanlong Zheng, Yanhao Zhang 等NeurIPS 2025 · 被引用 3 次
- Detailed 2D-3D Joint Representation for Human-Object InteractionYong-Lu Li, Xinpeng Liu, Han Lu, Shiyi Wang 等CVPR 2020
- Cascaded Human-Object Interaction RecognitionTianfei Zhou, Wenguan Wang, Siyuan Qi, Haibin Ling 等CVPR 2020
- KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human AnnotationsYang You, Yujing Lou, Chengkun Li, Zhoujun Cheng 等CVPR 2020
相关 Paper
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang 等ICCV 2025 · 被引用 21 次
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez 等CVPR 2021
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu 等ACM MM 2022 · 被引用 29 次
- MA-Bench: Towards Fine-grained Micro-Action UnderstandingKun Li, Jihao Gu, Fei Wang, Zhiliang Wu 等CVPR 2026 · 被引用 12 次
- Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and ChallengesTongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu 等CVPR 2024 · 被引用 23 次
