Linking the Characters: Video-oriented Social Graph Generation via Hierarchical-cumulative GCN
Shiwei Wu, Joya Chen, Tong Xu, Liyi Chen, Lingfei Wu, Yao Hu, Enhong Chen
Abstract
Recent years have witnessed the booming of online video platforms. Along this line, a graph to illustrate social relation among characters has been long expected to not only benefit the audiences for better understanding the story, but also support the fine-grained video analysis task in a semantic way. Unfortunately, though we humans could easily infer the social relations among characters, it is still an extremely challenging task for intelligent systems to automatically capture the social relation by absorbing multi-modal cues. Besides, they fail to describe the relations among multiple characters in a graph-generation perspective. To that end, inspired by the human inference ability on social relationship, we propose a novel Hierarchical- Cumulative Graph Convolutional Network (HC-GCN) to generate the social relation graph for multiple characters in the video. Specifically, we first integrate the short-term multi-modal cues, including visual, textual and audio information, to generate the frame-level graphs for part of characters via multimodal graph convolution technique. While dealing with the video-level aggregation task, we design an end-to-end framework to aggregate all frame-level subgraphs along the temporal trajectory, which results in a global video-level social graph with various social relationships among multiple characters. Extensive validations on two real-world large-scale datasets demonstrate the effectiveness of our proposed method compared with SOTA baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge GraphsLiyi Chen, Panrong Tong, Zhongming Jin, Ying Sun et al.NeurIPS 2024 · 160 citations
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang et al.NeurIPS 2024 · 78 citations
- Tackling Uncertain Correspondences for Multi-Modal Entity AlignmentLiyi Chen, Ying Sun, Shengzhe Zhang, Yuyang Ye et al.NeurIPS 2024 · 20 citations
- Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionHaorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang et al.EMNLP 2025 · 1 citation
Related papers
- Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long VideosHaorui Wang, Yibo Hu, Yangfu Zhu, Jinsheng Qi et al.ACM MM 2023 · 5 citations
- Hierarchical Cross-Modal Graph Consistency Learning for Video-Text RetrievalWeike Jin, Zhou Zhao, Pengcheng Zhang, Jieming Zhu et al.SIGIR 2021 · 49 citations
- Multi-Modal Multi-Action Video RecognitionZhensheng Shi, Ju Liang, Qianqian Li, Haiyong Zheng et al.ICCV 2021 · 11 citations
- Cause and Effect: Video Social Relationship Recognition from Causal PerspectiveYuxuan Zhang, Bo Wang, Yu Du, Yangfu Zhu et al.ACM MM 2025
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi et al.ICCV 2021 · 28 citations
