MPCHAT: Towards Multimodal Persona-Grounded Conversation
Jaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee Kim
摘要
In order to build self-consistent personalized dialogue agents, previous research has mostly focused on textual persona that delivers personal facts or personalities. However, to fully describe the multi-faceted nature of persona, image modality can help better reveal the speaker's personal characteristics and experiences in episodic memory (Rubin et al., 2003; Conway, 2009) . In this work, we extend persona-based dialogue to the multimodal domain and make two main contributions. First, we present the first multimodal persona-based dialogue dataset named MPCHAT, which extends persona with both text and images to contain episodic memories. Second, we empirically show that incorporating multimodal persona, as measured by three proposed multimodal persona-grounded dialogue tasks (i.e., next response prediction, grounding persona prediction, and speaker identification), leads to statistically significant performance improvements across all tasks. Thus, our work highlights that multimodal persona is crucial for improving multimodal dialogue comprehension, and our MPCHAT serves as a high-quality resource for this research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal 等ACL 2024 · 被引用 30 次
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park 等EMNLP 2023 · 被引用 5 次
- Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic InteractionsJihyoung Jang, Minwook Bae, Minji Kim, Dilek Hakkani-Tür 等ACL 2025
- Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text UpdatesJaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee KimACL 2025
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 被引用 316 次
相关 Paper
- Multimodal Persona Based Generation of Comic DialogsHarsh Agrawal, Aditya Mishra, Manish Gupta, MausamACL 2023 · 被引用 8 次
- Detecting Speaker Personas from Conversational TextsJia-Chen Gu, Zhen-Hua Ling, Yu Wu, Quan Liu 等EMNLP 2021
- Partner Matters! An Empirical Study on Fusing Personas for Personalized Response Selection in Retrieval-Based ChatbotsJia-Chen Gu, Hui Liu, Zhen-Hua Ling, Quan Liu 等SIGIR 2021 · 被引用 19 次
- PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueDongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang 等EMNLP 2025
- You Impress Me: Dialogue Generation via Mutual Persona PerceptionQian Liu, Yihong Chen, Bei Chen, Jian-Guang Lou 等ACL 2020 · 被引用 144 次
