Spatial-temporal Concept based Explanation of 3D ConvNets
Ying Ji, Yu Wang, Jien Kato
Abstract
Convolutional neural networks (CNNs) have shown remarkable performance on various tasks. Despite its widespread adoption, the decision procedure of the network still lacks transparency and interpretability, making it difficult to enhance the performance further. Hence, there has been considerable interest in providing explanation and interpretability for CNNs over the last few years. Explainable artificial intelligence (XAI) investigates the relationship between input images or videos and output predictions. Recent studies have achieved outstanding success in explaining 2D image classification ConvNets. On the other hand, due to the high computation cost and complexity of video data, the explanation of 3D video recognition ConvNets is relatively less studied. And none of them are able to produce a high-level explanation. In this paper, we propose a STCE (Spatial-temporal Concept-based Explanation) framework for interpreting 3D ConvNets. In our approach: (1) videos are represented with high-level supervoxels, similar supervoxels are clustered as a concept, which is straightforward for human to understand; and (2) the interpreting framework calculates a score for each concept, which reflects its significance in the ConvNet decision procedure. Experiments on diverse 3D ConvNets demonstrate that our method can identify global concepts with different importance levels, allowing us to investigate the impact of the concepts on a target task, such as action recognition, in-depth. The source codes are publicly available at https://github.com/yingji425/STCE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65a2f7ea-705c-4830-81f1-39ef8d5300f7Cited by top-tier papers3
- State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User UnderstandingDevleena Das, Sonia Chernova, Been KimNeurIPS 2023 · 33 citations
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action RecognitionJongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim et al.NeurIPS 2025 · 4 citations
- On the Variability of Concept Activation VectorsJulia Wenkmann, Damien GarreauICML 2026 · 3 citations
Builds on3
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Interpretable Neural Networks with Frank-Wolfe: Sparse Relevance Maps and Relevance OrderingsJan MacDonald, Mathieu Besançon, Sebastian PokuttaICML 2022 · 13 citations
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng et al.CVPR 2021
Related papers
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon et al.CVPR 2024
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang et al.ACM MM 2021 · 3 citations
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger et al.AAAI 2021 · 140 citations
- Interpretable 3D Neural Object Volumes for Robust Conceptual ReasoningNhi Pham, Artur Jesslen, Bernt Schiele, Adam Kortylewski et al.ICLR 2026 · 2 citations
- Towards Global Explanations of Convolutional Neural Networks With Concept AttributionWeibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao et al.CVPR 2020
