Cause and Effect: Video Social Relationship Recognition from Causal Perspective
Yuxuan Zhang, Bo Wang, Yu Du, Yangfu Zhu, Haorui Wang, Guangyao Su, Tao Zhou, Bin Wu
Abstract
Video social relation recognition is a fundamental task in video understanding, which is dedicated to the construction of multi-modal knowledge graphs. Previous work mainly focuses on multi-modal fusion and the construction of special character graphs. However, they often treat the global frame sequence equally, ignoring the influence of key frame sequence on relation recognition. Specifically, the key frame sequence that significantly reflect character relationships in a video tends to be sparse and short. At the same time, the key frames have not only temporal but also strong causal relationship. Therefore, we propose a novel Video Local Causal Frame (VLCF) model to explore the causal relationship between frames. Inspired by Granger causality theory, we estimate inter-frame causal relationships by comparing the predicted result frames with and without masking the premise frame. We then construct global connections between video frames. Multiple local causal frame sequences and global frame sequences are extracted to capture the key information and global information in the video. Extensive experiments conducted on the ViSR dataset and the MovieGraphs dataset demonstrate that the proposed model achieves state-of-the-art performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionHaorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang et al.EMNLP 2025 · 1 citation
- Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long VideosHaorui Wang, Yibo Hu, Yangfu Zhu, Jinsheng Qi et al.ACM MM 2023 · 5 citations
- Linking the Characters: Video-oriented Social Graph Generation via Hierarchical-cumulative GCNShiwei Wu, Joya Chen, Tong Xu, Liyi Chen et al.ACM MM 2021 · 26 citations
- Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceJiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin et al.ACM MM 2023 · 7 citations
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 6 citations
