Object-Relation Reasoning Graph for Action Recognition
Yangjun Ou, Li Mi, Zhenzhong Chen
Abstract
Action recognition is a challenging task since the attributes of objects as well as their relationships change constantly in the video. Existing methods mainly use object-level graphs or scene graphs to represent the dynamics of objects and relationships, but ignore modeling the fine-grained relationship transitions directly. In this paper, we propose an Object-Relation Reasoning Graph (OR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> G) for reasoning about action in videos. By combining an object-level graph (OG) and a relation-level graph (RG), the proposed OR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> G catches the attribute transitions of objects and reasons about the relationship transitions between objects simultaneously. In addition, a graph aggregating module (GAM) is investigated by applying the multi-head edge-to-node message passing operation. GAM feeds back the information from the relation node to the object node and enhances the coupling between the object-level graph and the relation-level graph. Experiments in video action recognition demonstrate the effectiveness of our approach when compared with the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10e22cad-436f-4fed-9e52-5253d9eba487Cited by top-tier papers5
- : A Visual Analytics Approach for Interactive Video ProgrammingJianben He, Xingbo Wang, Kamkwai Wong, Xijie Huang et al.IEEE VIS 2023 · 17 citations
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang et al.CVPR 2024 · 3 citations
- Punching Bag vs. Punching Person: Motion Transferability in VideosRaiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran et al.ICCV 2025 · 1 citation
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang et al.NeurIPS 2025 · 1 citation
- Action Detail Matters: Refining Video Recognition with Local Action QueriesMengmeng Wang, Zeyi Huang, Xiangjie Kong, Guojiang Shen et al.CVPR 2025
Builds on6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu et al.CVPR 2020
- MoViNets: Mobile Video Networks for Efficient Video RecognitionDan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang et al.CVPR 2021
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
Related papers
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 57 citations
- Temporal Relational Modeling with Self-Supervision for Action SegmentationDong Wang, Di Hu, Xingjian Li, Dejing DouAAAI 2021 · 63 citations
- Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionHaorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang et al.EMNLP 2025 · 1 citation
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- Edge-Centric Relational Reasoning for 3D Scene Graph PredictionYanni Ma, Hao Liu, Yulan Guo, Theo Gevers et al.AAAI 2026
