Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding
Yueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang, Qun Liu, Dongyan Zhao
Abstract
Multi-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal conversations, MMC requires stronger character-centered understanding abilities as there are many interlocutors appearing in both the visual and textual context. To facilitate the study of this problem, we present Friends-MMC in this paper, an MMC dataset that contains 24,000+ unique utterances paired with video context. To explore the character-centered understanding of the dialogue, we also annotate the speaker of each utterance, the names and bounding bboxes of faces that appear in the video. Based on this Friends-MMC dataset, we further study two fundamental MMC tasks: conversation speaker identification and conversation response prediction, both of which have the multi-party nature with the video or image as visual context. For conversation speaker identification, we demonstrate the inefficiencies of existing methods such as pre-trained models, and propose a simple yet effective baseline method that leverages an optimization solver to utilize the context of two modalities to achieve better performance. For conversation response prediction, we fine-tune generative dialogue models on Friend-MMC, and analyze the benefits of speaker information. The code and dataset will be publicly available, and thus we call for more attention on modelling speaker information when understanding conversations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e639a3e6-006c-4a8d-8583-0ecb2040d1faCited by top-tier papers1
Ask how each one uses itBuilds on13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou et al.ICLR 2024 · 303 citations
Related papers
- A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party ConversationsWenjie Zheng, Jianfei Yu, Rui Xia, Shijin WangACL 2023 · 38 citations
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla et al.CVPR 2024
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingJia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu et al.ACL 2021
- Seeing Conversations: Communication Context Identification in Egocentric VideoTobias Dorszewski, Jens HjortkjærCVPR 2026
