FL-MSRE: A Few-Shot Learning based Approach to Multimodal Social Relation Extraction
Hai Wan, Manrong Zhang, Jianfeng Du, Ziling Huang, Yufei Yang, Jeff Z. Pan
Abstract
Social relation extraction (SRE for short), which aims to infer the social relation between two people in daily life, has been demonstrated to be of great value in reality. Existing methods for SRE consider extracting social relation only from unimodal information such as text or image, ignoring the high coupling of multimodal information. Moreover, previous studies overlook the serious unbalance distribution on social relations. To address these issues, this paper proposes FL-MSRE, a few-shot learning based approach to extracting social relations from both texts and face images. Considering the lack of multimodal social relation datasets, this paper also presents three multimodal datasets annotated from four classical masterpieces and corresponding TV series. Inspired by the success of BERT, we propose a strong BERT based baseline to extract social relation from text only. FL-MSRE is empirically shown to outperform the baseline significantly. This demonstrates that using face images benefits text-based SRE. Further experiments also show that using two faces from different images achieves similar performance as from the same image. This means that FL-MSRE is suitable for a wide range of SRE applications where the faces of two people can only be collected from different images. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- Multimodal Relation Extraction via a Mixture of Hierarchical Visual Context LearnersXiyang Liu, Chunming Hu, Richong Zhang, Kai Sun et al.WWW 2024 · 21 citations
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information ExtractionBaohang Zhou, Ying Zhang, Yu Zhao, Xuhui Sui et al.WWW 2025 · 5 citations
- Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information ExtractionBaohang Zhou, Kehui Song, Rize Jin, Yu Zhao et al.WWW 2026
Builds on2
Related papers
- Linking People across Text and Images Based on Social Relation ReasoningYang Lei, Peizhi Zhao, Pijian Li, Yi Cai et al.AAAI 2023
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai et al.ACM MM 2021 · 134 citations
- Few-Shot Joint Multimodal Entity-Relation Extraction via Knowledge-Enhanced Cross-modal Prompt ModelLi Yuan, Yi Cai, Junsheng HuangACM MM 2024 · 9 citations
- Towards relation extraction from speechTongtong Wu, Guitao Wang, Jinming Zhao, Zhaoran Liu et al.EMNLP 2022 · 6 citations
- Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation TaggingLi Yuan, Yi Cai, Jin Wang, Qing LiAAAI 2023 · 89 citations
