Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction Recognition
Ning Wang, Guangming Zhu, Liang Zhang, Peiyi Shen, Hongsheng Li, Cong Hua
摘要
For a given video-based Human-Object Interaction scene, modeling the spatio-temporal relationship between humans and objects is the important cue to understand the contextual information presented in the video. With the efficient spatio-temporal relationship modeling, it is possible not only to uncover contextual information in each frame, but to directly capture inter-frame dependencies as well. Capturing the position changes of human and objects over the spatio-temporal dimension is more critical when significant changes in the appearance features may not occur over time. When utilizing appearance features, the spatial location and the semantic information are also the key to improve the video-based Human-Object Interaction recognition performance. In this paper, Spatio-Temporal Interaction Graph Parsing Networks (STIGPN) are constructed, which encode the videos with a graph composed of human and object nodes. These nodes are connected by two types of relations: (i) intra-frame relations: modeling the interactions between human and the interacted objects within each frame. (ii) inter-frame relations: capturing the long range dependencies between human and the interacted objects across frame. With the graph, STIGPN learn spatio-temporal features directly from the whole video-based Human-Object Interaction scenes. Multi-modal features and a multi-stream fusion strategy are used to enhance the reasoning capability of STIGPN. Two Human-Object Interaction video datasets, including CAD-120 and Something-Else, are used to evaluate the proposed architectures, and the state-of-the-art performance demonstrates the superiority of STIGPN. Code for STIGPN is available at https://github.com/GuangmingZhu/STIGPN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Video-based Human-Object Interaction Detection from Tubelet TokensDanyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai 等NeurIPS 2022 · 被引用 24 次
- Open Set Video HOI detection from Action-centric Chain-of-Look PromptingNan Xi, Jingjing Meng, Junsong YuanICCV 2023 · 被引用 10 次
- MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang 等ACL 2026 · 被引用 10 次
- Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion PredictionWanying Zhang, Shen Zhao, Fanyang Meng, Songtao Wu 等ACM MM 2023 · 被引用 8 次
- Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image EditingChangHee Yang, Chanhee Kang, Kyeongbo Kong, Hanni Oh 等CVPR 2024
它引用的顶会 Paper4
- Reasoning About Human-Object Interactions Through Dual Attention NetworksTete Xiao, Quanfu Fan, Danny Gutfreund, Mathew Monfort 等ICCV 2019 · 被引用 36 次
- LIGHTEN: Learning Interactions with Graph and Hierarchical TEmporal Networks for HOI in videosSai Praneeth Reddy Sunkesula, Rishabh Dabral, Ganesh RamakrishnanACM MM 2020 · 被引用 36 次
- Context Aware Graph Convolution for Skeleton-Based Action RecognitionXikun Zhang, Chang Xu, Dacheng TaoCVPR 2020
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu 等CVPR 2020
相关 Paper
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 被引用 47 次
- Relation Parsing Neural Network for Human-Object Interaction DetectionPenghao Zhou, Mingmin ChiICCV 2019 · 被引用 155 次
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang 等NeurIPS 2025 · 被引用 1 次
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 被引用 5 次
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 被引用 57 次
