MPCHAT: Towards Multimodal Persona-Grounded Conversation
Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim
Abstract
In order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities. However, to fully describe the multi-faceted nature of persona, image modality can help better reveal the speaker's personal characteristics and experiences in episodic memory (Rubin et al., 2003; Conway, 2009) . In this work, we extend persona-based dialogue to the multimodal domain and make two main contributions. First, we present the first multimodal persona-based dialogue dataset named MPCHAT, which extends persona with both text and images to contain episodic memories. Second, we empirically show that incorporating multimodal persona, as measured by three proposed multimodal persona-grounded dialogue tasks (i.e., next response prediction, grounding persona prediction, and speaker identification), leads to statistically significant performance improvements across all tasks. Thus, our work highlights that multimodal persona is crucial for improving multimodal dialogue comprehension, and our MPCHAT serves as a high-quality resource for this research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 930dee06-631d-4883-a991-5392f5338169Cited by top-tier papers4
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal et al.ACL 2024 · 30 citations
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park et al.EMNLP 2023 · 5 citations
- Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic InteractionsJihyoung Jang, Minwook Bae, Minji Kim, Dilek Hakkani-Tür et al.ACL 2025
- Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text UpdatesJaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee KimACL 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
Related papers
- Multimodal Persona Based Generation of Comic DialogsHarsh Agrawal, Aditya Mishra, Manish Gupta, MausamACL 2023 · 8 citations
- Detecting Speaker Personas from Conversational TextsJia-Chen Gu, Zhen-Hua Ling, Yu Wu, Quan Liu et al.EMNLP 2021
- Partner Matters! An Empirical Study on Fusing Personas for Personalized Response Selection in Retrieval-Based ChatbotsJia-Chen Gu, Hui Liu, Zhen-Hua Ling, Quan Liu et al.SIGIR 2021 · 19 citations
- PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueDongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang et al.EMNLP 2025
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou et al.ACL 2020 · 144 citations
