Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long Videos
Haorui Wang, Yibo Hu, Yangfu Zhu, Jinsheng Qi, Bin Wu
摘要
Social Relation Recognition is an important part of Video Understanding, providing insights into the information that videos convey. Most previous works mainly focused on graph generation for characters, instead of edges which are more suitable for relation modelling. Furthermore, previous methods tend to recognize social relations for single frames or short video clips within their receptive fields, neglecting the importance of continuous reasoning throughout the entire video. To tackle these challenges, we propose a novel Shifted GCN-GAT and Cumulative-Transformer framework, named SGCAT-CT. The overall architecture consists of an SGCAT module for shifted graph operations on novel relation graphs and a CT module for temporal processing with memory. SGCAT-CT conducts continuous recognition of social relations and memorizes information from as early as the beginning of a long video. Experiments conducted on several video datasets demonstrate encouraging performance on long videos. Our code will be released at https://github.com/HarryWgCN/SGCAT-CT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment OptimizationWanhua Li, Zibin Meng, Jiawei Zhou, Donglai Wei 等NeurIPS 2024 · 被引用 20 次
- Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionHaorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang 等EMNLP 2025 · 被引用 1 次
相关 Paper
- Linking the Characters: Video-oriented Social Graph Generation via Hierarchical-cumulative GCNShiwei Wu, Joya Chen, Tong Xu, Liyi Chen 等ACM MM 2021 · 被引用 26 次
- Temporal Relational Modeling with Self-Supervision for Action SegmentationDong Wang, Di Hu, Xingjian Li, Dejing DouAAAI 2021 · 被引用 63 次
- Cause and Effect: Video Social Relationship Recognition from Causal PerspectiveYuxuan Zhang, Bo Wang, Yu Du, Yangfu Zhu 等ACM MM 2025
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 被引用 4 次
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 被引用 37 次
