Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation Chains
Hua Hua, Dongxu Li, Ruiqi Li, Peng Zhang, Jochen Renz, Anthony G. Cohn
摘要
In order to be trusted by humans, Artificial Intelligence agents should be able to describe rationales behind their decisions. One such application is human action recognition in critical or sensitive scenarios, where trustworthy and explainable action recognizers are expected. For example, reliable pedestrian action recognition is essential for self-driving cars and explanations for real-time decision making are critical for investigations if an accident happens. In this regard, learning-based approaches, despite their popularity and accuracy, are disadvantageous due to their limited interpretability.
This paper presents a novel neuro-symbolic approach that recognizes actions from videos with human-understandable explanations. Specifically, we first propose to represent videos symbolically by qualitative spatial relations between objects called qualitative spatial object relation chains. We further develop a neural saliency estimator to capture the correlation between such object relation chains and the occurrence of actions. Given an unseen video, this neural saliency estimator is able to tell which object relation chains are more important for the action recognized. We evaluate our approach on two real-life video datasets, with respect to recognition accuracy and the quality of generated action explanations. Experiments show that our approach achieves superior performance on both aspects to previous symbolic approaches, thus facilitating trustworthy intelligent decision making. Our approach can be used to augment state-of-the-art learning approaches with explainabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang 等CVPR 2024 · 被引用 3 次
- Axiomatizability of Alexandrov Dynamic Topological LogicNiels C. Vooijs, David Fernández-DuqueLICS 2026
它引用的顶会 Paper7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang 等NeurIPS 2020 · 被引用 171 次
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang 等CVPR 2021
相关 Paper
- Generating Explanations for Embodied Action Decision from Visual ObservationXiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang 等ACM MM 2023 · 被引用 3 次
- Spatial-temporal Concept based Explanation of 3D ConvNetsYing Ji, Yu Wang, Jien KatoCVPR 2023
- An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue GenerationShiquan Yang, Rui Zhang, Sarah M. Erfani, Jey Han LauACL 2022 · 被引用 17 次
- Complex Video Action Reasoning via Learnable Markov Logic NetworkYang Jin, Linchao Zhu, Yadong MuCVPR 2022 · 被引用 13 次
- Motion Question Answering via Modular Motion ProgramsMark Endo, Joy Hsu, Jiaman Li, Jiajun WuICML 2023 · 被引用 28 次
