MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents
Yanqi Dai, Huanran Hu, Lei Wang, Shengjie Jin, Xu Chen, Zhiwu Lu
摘要
Recently, Role-Playing Agents (RPAs) have garnered increasing attention for their potential to deliver emotional value and facilitate sociological research. However, existing studies are primarily confined to the textual modality, unable to simulate humans' multimodal perceptual capabilities. To bridge this gap, we introduce the concept of Multimodal Role-Playing Agents (MRPAs), and propose a comprehensive framework, MMRole, for their development and evaluation, which comprises a personalized multimodal dataset and a robust evaluation approach. Specifically, we construct a large-scale, high-quality dataset, MMRole-Data, consisting of 85 characters, 11K images, and 14K single or multi-turn dialogues. Additionally, we present a robust evaluation approach, MMRole-Eval, encompassing eight metrics across three dimensions, where a reward model is designed to score MRPAs with the constructed ground-truth data for comparison. Moreover, we develop the first specialized MRPA, MMRole-Agent. Extensive evaluation results demonstrate the improved performance of MMRole-Agent and highlight the primary challenges in developing MRPAs, emphasizing the need for enhanced multimodal understanding and role-playing consistency. The data, code, and models are all available at https://github.com/YanqiDai/MMRole.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing AgentsYuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu 等ICLR 2026 · 被引用 2 次
- Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing AgentsXueqiao Zhang, Chao Zhang, Jingtao Xu, Yifan Zhu 等EMNLP 2025 · 被引用 2 次
- DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory GraphZhihao Xiao, Mengting Li, Xintao Wang, Linfeng Li 等KDD 2026 · 被引用 1 次
- CoSER: Coordinating LLM-Based Persona Simulation of Established RolesXintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan 等ICML 2025
- MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion UnderstandingZhicheng Zhang, Wuyou Xia, Chenxi Zhao, Zhou Yan 等ICML 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent EvaluationQuan Tu, Shilong Fan, Zihang Tian, Tianhao Shen 等ACL 2024
- ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character ChatLanlan Qiu, Sophia Xiao Pu, Yeqi Feng, Wenchang Gao 等ACL 2026
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin 等EMNLP 2024 · 被引用 2 次
- OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality InteractionHaonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu 等ACL 2025 · 被引用 10 次
- RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing AgentsPinyi Zhang, Siyu An, Lingfeng Qiao, Yifei Yu 等ACL 2025 · 被引用 4 次
