Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim, Jinwoo Choi
摘要
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods-based on saliency-produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature-intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets-KTH, Penn Action, HAA500, and UCF-101-demonstrate that DANCE significantly improves explanation clarity with competitive performance. We validate the superior interpretability of DANCE through a user study. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis. Our project page is available at https://jong980812.github.io/DANCE/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
相关 Paper
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon 等CVPR 2024
- Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yingya Zhang 等ICCV 2023 · 被引用 40 次
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo 等NeurIPS 2025
- Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition *Xiaodan Hu, Narendra AhujaICCV 2021 · 被引用 28 次
- Spatial-temporal Concept based Explanation of 3D ConvNetsYing Ji, Yu Wang, Jien KatoCVPR 2023
