DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and Agent
Dingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li, Chengfu Tang, Dongxu Zhang, Zhenkui Wang, Xudong Wang, Song Liu
摘要
Recent years have witnessed a profound evolution in the abilities of Large Language Model, which has significantly boosted the proliferation of role-playing agents and platforms. Nonetheless, there is a conspicuous absence of systematic and comprehensive evaluations of role-playing abilities which are truly aligned with users' interaction scenarios in real-world. To address this gap, we have devised DMT-RoleBench, a benchmark designed to evaluate the role-playing abilities of large language models and agents based on dynamic multi-turn dialogues. Compared with existed role-playing benchmarks, DMT-RoleBench boasts several principal advantages: (1) It contains a more diverse role types and system prompts of different formats. (2) We propose an innovative evaluation paradigm to assess role-playing abilities based on dynamically generating multi-turn dialogues constrained by specific evaluation intents and topics, which is well aligned with users' interaction scenarios in real-world. (3) We define a three-tiered metric system and provide DMT-RM, which is a reward model aligned with human annotations, to annotate the dialogues. And we propose DMT-Score to calculate the final scores based on the annotated dialogues. Our experiments and analysis of leading models equipped with role-playing abilities have demonstrated the effectiveness of DMT-RoleBench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language ModelsZhouxing Tan, Ruochong Xiong, Yulong Wan, Jinlong Ma 等AAAI 2026
- ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character ChatLanlan Qiu, Sophia Xiao Pu, Yeqi Feng, Wenchang Gao 等ACL 2026
它引用的顶会 Paper5
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 被引用 97 次
- Personalized Dialogue Generation with Persona-Adaptive AttentionQiushi Huang, Yu Zhang, Tom Ko, Xubo Liu 等AAAI 2023 · 被引用 41 次
- Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-AlignmentKeming Lu, Bowen Yu, Chang Zhou, Jingren ZhouACL 2024 · 被引用 16 次
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang 等ACL 2021
- CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent EvaluationQuan Tu, Shilong Fan, Zihang Tian, Tianhao Shen 等ACL 2024
相关 Paper
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
- Crab: A Novel Configurable Role-Playing LLM with Assessing BenchmarkKai He, Yucheng Huang, Wenqing Wang, Delong Ran 等ACL 2025
- LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language ModelsMarwa Abdulhai, Isadora White, Charlie Victor Snell, Charles Sun 等ICML 2025
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation MetricsChenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang 等ACL 2026 · 被引用 10 次
- clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational AgentsKranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira 等EMNLP 2023 · 被引用 6 次
