Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation Chains
Hua Hua, Dongxu Li, Ruiqi Li, Peng Zhang, Jochen Renz, Anthony G. Cohn
Abstract
In order to be trusted by humans, Artificial Intelligence agents should be able to describe rationales behind their decisions. One such application is human action recognition in critical or sensitive scenarios, where trustworthy and explainable action recognizers are expected. For example, reliable pedestrian action recognition is essential for self-driving cars and explanations for real-time decision making are critical for investigations if an accident happens. In this regard, learning-based approaches, despite their popularity and accuracy, are disadvantageous due to their limited interpretability.
This paper presents a novel neuro-symbolic approach that recognizes actions from videos with human-understandable explanations. Specifically, we first propose to represent videos symbolically by qualitative spatial relations between objects called qualitative spatial object relation chains. We further develop a neural saliency estimator to capture the correlation between such object relation chains and the occurrence of actions. Given an unseen video, this neural saliency estimator is able to tell which object relation chains are more important for the action recognized. We evaluate our approach on two real-life video datasets, with respect to recognition accuracy and the quality of generated action explanations. Experiments show that our approach achieves superior performance on both aspects to previous symbolic approaches, thus facilitating trustworthy intelligent decision making. Our approach can be used to augment state-of-the-art learning approaches with explainabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f9f6bc8-8c46-493e-934a-e1577a2c2f77Cited by top-tier papers2
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang et al.CVPR 2024 · 3 citations
- Axiomatizability of Alexandrov Dynamic Topological LogicNiels C. Vooijs, David Fernández-DuqueLICS 2026
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang et al.CVPR 2021
Related papers
- Generating Explanations for Embodied Action Decision from Visual ObservationXiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang et al.ACM MM 2023 · 3 citations
- Spatial-temporal Concept based Explanation of 3D ConvNetsYing Ji, Yu Wang, Jien KatoCVPR 2023
- An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue GenerationShiquan Yang, Rui Zhang, Sarah M. Erfani, Jey Han LauACL 2022 · 17 citations
- Complex Video Action Reasoning via Learnable Markov Logic NetworkYang Jin, Linchao Zhu, Yadong MuCVPR 2022 · 13 citations
- Motion Question Answering via Modular Motion ProgramsMark Endo, Joy Hsu, Jiaman Li, Jiajun WuICML 2023 · 28 citations
