Complex Video Action Reasoning via Learnable Markov Logic Network
Yang Jin, Linchao Zhu, Yadong Mu
Abstract
Profiting from the advance of deep convolutional networks, current state-of-the-art video action recognition models have achieved remarkable progress. Nevertheless, most of existing models suffer from low interpretability of the predicted actions. Inspired by the observation that temporally-configured human-object interactions often serve as a key indicator of many actions, this work crafts an action reasoning framework that performs Markov Logic Network (MLN) based probabilistic logical inference. Crucially, we propose to encode an action by first-order logical rules that correspond to the temporal changes of visual relationships in videos. The main contributions of this work are two-fold: 1) Different from existing black-box models, the proposed model simultaneously implements the localization of temporal boundaries and the recognition of action categories by grounding the logical rules of MLN in videos. The weight associated with each such rule further provides an estimate of confidence. These collectively make our model more explainable and robust. 2) Instead of using hand-crafted logical rules in conventional MLN, we develop a data-driven instantiation of the MLN. In specific, a hybrid learning scheme is proposed. It combines MLN's weight learning and reinforcement learning, using the former's results as a self-critic for guiding the latter's training. Additionally, by treating actions as logical predicates, the proposed framework can also be integrated with deep models for further performance boost. Comprehensive experiments on two complex video action datasets (Charades & CAD-120) clearly demonstrate the effectiveness and explainability of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen et al.NeurIPS 2024 · 36 citations
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang et al.CVPR 2024 · 3 citations
- TRACE: Temporal Grounding Video LLM via Causal Event ModelingYongxin Guo, Jingyu Liu, Mingda Li, Qingbin Liu et al.ICLR 2025
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- DynamoNet: Dynamic Action and Motion NetworkAli Diba, Vivek Sharma, Luc Van Gool, Rainer StiefelhagenICCV 2019 · 123 citations
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Graph-Based High-Order Relation Modeling for Long-Term Action RecognitionJiaming Zhou, Kun-Yu Lin, Haoxin Li, Wei-Shi ZhengCVPR 2021
Related papers
- A Probabilistic Graphical Model Based on Neural-symbolic Reasoning for Visual Relationship DetectionDongran Yu, Bo Yang, Qianhao Wei, Anchen Li et al.CVPR 2022 · 18 citations
- Improving Out-of-Distribution Detection with Markov Logic NetworksKonstantin Kirchheim, Frank OrtmeierICML 2025
- ExCAR: Event Graph Knowledge Enhanced Explainable Causal ReasoningLi Du, Xiao Ding, Kai Xiong, Ting Liu et al.ACL 2021
- Zero-shot Compositional Action Recognition with Neural Logic ConstraintsGefan Ye, Lin Li, Kexin Li, Jun Xiao et al.ACM MM 2025 · 1 citation
- Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question AnsweringZihan Song, Xin Wang, Zi Qian, Hong Chen et al.ICML 2025
