Object-Relation Reasoning Graph for Action Recognition
Yangjun Ou, Li Mi, Zhenzhong Chen
摘要
Action recognition is a challenging task since the attributes of objects as well as their relationships change constantly in the video. Existing methods mainly use object-level graphs or scene graphs to represent the dynamics of objects and relationships, but ignore modeling the fine-grained relationship transitions directly. In this paper, we propose an Object-Relation Reasoning Graph (OR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> G) for reasoning about action in videos. By combining an object-level graph (OG) and a relation-level graph (RG), the proposed OR <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> G catches the attribute transitions of objects and reasons about the relationship transitions between objects simultaneously. In addition, a graph aggregating module (GAM) is investigated by applying the multi-head edge-to-node message passing operation. GAM feeds back the information from the relation node to the object node and enhances the coupling between the object-level graph and the relation-level graph. Experiments in video action recognition demonstrate the effectiveness of our approach when compared with the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- : A Visual Analytics Approach for Interactive Video ProgrammingJianben He, Xingbo Wang, Kamkwai Wong, Xijie Huang 等IEEE VIS 2023 · 被引用 17 次
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang 等CVPR 2024 · 被引用 3 次
- Punching Bag vs. Punching Person: Motion Transferability in VideosRaiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran 等ICCV 2025 · 被引用 1 次
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang 等NeurIPS 2025 · 被引用 1 次
- Action Detail Matters: Refining Video Recognition with Local Action QueriesMengmeng Wang, Zeyi Huang, Xiangjie Kong, Guojiang Shen 等CVPR 2025
它引用的顶会 Paper6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu 等CVPR 2020
- MoViNets: Mobile Video Networks for Efficient Video RecognitionDan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang 等CVPR 2021
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
相关 Paper
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 被引用 57 次
- Temporal Relational Modeling with Self-Supervision for Action SegmentationDong Wang, Di Hu, Xingjian Li, Dejing DouAAAI 2021 · 被引用 63 次
- Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionHaorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang 等EMNLP 2025 · 被引用 1 次
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 被引用 47 次
- Edge-Centric Relational Reasoning for 3D Scene Graph PredictionYanni Ma, Hao Liu, Yulan Guo, Theo Gevers 等AAAI 2026
