ACL2026
LongMP-Bench: A Benchmark for Multimodal Persona Understanding in Long-Term Dialogues
Zhuoqun Li, Zhaopei Huang, Wenxuan Wang, Qin Jin
摘要
Understanding multimodal user personas over long-term dialogues is essential for personalized and human-like dialogue systems. In realistic interactions, user personas evolve over time and are expressed through both language and visual cues. However, existing benchmarks provide limited support for evaluating such dynamic and multimodal persona understanding, due to shallow persona coverage, weak visual consistency, and static settings. We introduce LongMP-Bench, a benchmark for evaluating models' ability to understand, track, and utilize evolving multimodal user personas in long-term dialogues. We propose a scalable, multi-step data construction pipeline to synthesize extended multimodal interactions, followed by human refinement. The resulting dataset contains long-term conversations from 150 distinct users, each maintaining visual identity consistency while exhibiting progressive persona evolution. Based on LongMP-Bench, we define evaluation tasks for persona tracking, multimodal reasoning, and personalized response generation. Extensive experiments show that current multimodal large language models struggle with long-term persona consistency, persona shifts, and effective multimodal integration. Our data and code are available at https://github.com/skspass/ LongMP-Bench .