Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
Jongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim, Jinwoo Choi
Abstract
Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods-based on saliency-produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature-intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets-KTH, Penn Action, HAA500, and UCF-101-demonstrate that DANCE significantly improves explanation clarity with competitive performance. We validate the superior interpretability of DANCE through a user study. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis. Our project page is available at https://jong980812.github.io/DANCE/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68ec84b2-fe4c-4b6c-a8d3-d725480425acBuilds on27
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
Related papers
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon et al.CVPR 2024
- Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningZhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yingya Zhang et al.ICCV 2023 · 40 citations
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo et al.NeurIPS 2025
- Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition *Xiaodan Hu, Narendra AhujaICCV 2021 · 28 citations
- Spatial-temporal Concept based Explanation of 3D ConvNetsYing Ji, Yu Wang, Jien KatoCVPR 2023
