Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation Recognition
Haorui Wang, Zheng Wang, Yuxuan Zhang, Bo Wang, Bin Wu
Abstract
Recent years have witnessed remarkable advances in Large Language Models (LLMs). However, in the task of social relation recognition, Large Language Models (LLMs) encounter significant challenges due to their reliance on sequential training data, which inherently restricts their capacity to effectively model complex graph-structured relationships. To address this limitation, we propose a novel low-coupling method synergizing multimodal temporal Knowledge Graphs and Large Language Models (mtKG-LLM) for social relation reasoning. Specifically, we extract multimodal information from the videos and model the social networks as spatial Knowledge Graphs (KGs) for each scene. Temporal KGs are constructed based on spatial KGs and updated along the timeline for long-term reasoning. Subsequently, we retrieve multi-scale information from the graph-structured knowledge for LLMs to recognize the underlying social relation. Extensive experiments demonstrate that our method has achieved state-ofthe-art performance in social relation recognition. Furthermore, our framework exhibits effectiveness in bridging the gap between KGs and LLMs. We release our code at https: //github.com/HarryWgCN/mtKG-LLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 202d749e-58e1-4098-9fdc-576a3d2f34edBuilds on14
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi et al.SIGIR 2024 · 182 citations
- StructGPT: A General Framework for Large Language Model to Reason over Structured DataJinhao Jiang, Kun Zhou, Zican Dong, Keming Ye et al.EMNLP 2023 · 173 citations
- Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph ReasoningJiapu Wang, Kai Sun, Linhao Luo, Wei Wei et al.NeurIPS 2024 · 82 citations
- Enhanced Story Comprehension for Large Language Models through Dynamic Document-Based Knowledge GraphsBerkeley R. Andrus, Yeganeh Nasiri, Shilong Cui, Benjamin Cullen et al.AAAI 2022 · 46 citations
Related papers
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long VideosHaorui Wang, Yibo Hu, Yangfu Zhu, Jinsheng Qi et al.ACM MM 2023 · 5 citations
- SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and AudioTong Lu, Zhichun Wang, Yaoyu Zhou, Yiming Guan et al.AAAI 2026
- STK-Adapter: Incorporating Evolving Graph and Event Chain for Temporal Knowledge Graph ExtrapolationShuyuan Zhao, Wei Chen, Weijie Zhang, Xinrui Hou et al.ACL 2026
- Cause and Effect: Video Social Relationship Recognition from Causal PerspectiveYuxuan Zhang, Bo Wang, Yu Du, Yangfu Zhu et al.ACM MM 2025
