CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, Rui Yan
Abstract
Recently, the advent of large language models 001 (LLMs) has revolutionized generative agents. 002 Among them, Role-Playing Conversational 003 Agents (RPCAs) attract considerable atten-004 tion due to their ability to emotionally engage 005 users. However, the absence of a compre-006 hensive benchmark impedes progress in this 007 field. To bridge this gap, we introduce Char-008 acterEval, a Chinese benchmark for compre-009 hensive RPCA assessment, complemented by a 010 tailored high-quality dataset. The dataset com-011 prises 1,785 multi-turn role-playing dialogues, 012 encompassing 11,376 examples and featuring 013 77 characters derived from Chinese novels and 014 scripts. It was carefully constructed, beginning 015 with initial dialogue extraction via GPT-4, fol-016 lowed by rigorous human-led quality control, 017 and enhanced with in-depth character profiles 018 sourced from Baidu Baike. CharacterEval em-019 ploys a multifaceted evaluation approach, en-020 compassing thirteen targeted metrics on four 021 dimensions. To facilitate the convenient eval-022 uation for these subjective metrics in Charac-023 terEval, we further developed CharacterRM, a 024 role-playing reward model based on human an-025 notations, which has a higher correlation with 026 human judgment compared to GPT-4. Compre-027 hensive experiments on CharacterEval demon-028 strate that Chinese LLMs exhibit more promis-029 ing capabilities than GPT-4 in Chinese role-030 playing conversation 1 . 031 1 Introduction 032 The development of large language models (LLMs) 033 has marked the beginning of a new era in conversa-034
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers47
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou et al.ACL 2025 · 24 citations
- Beyond Dialogue: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language ModelYeyong Yu, Runsheng Yu, Haojie Wei, Zhanqiu Zhang et al.ACL 2025 · 13 citations
- OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality InteractionHaonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu et al.ACL 2025 · 10 citations
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsMuyu He, Anand Kumar, Soumyadeep Bakshi, James Zou et al.ACL 2026 · 8 citations
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li et al.AAAI 2025 · 6 citations
Builds on3
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen et al.EMNLP 2020 · 78 citations
- Zero-Resource Knowledge-Grounded Dialogue GenerationLinxiao Li, Can Xu, Wei Wu, Yufan Zhao et al.NeurIPS 2020 · 75 citations
Related papers
- CharacterBench: Benchmarking Character Customization of Large Language ModelsJinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi et al.AAAI 2025 · 7 citations
- Crab: A Novel Configurable Role-Playing LLM with Assessing BenchmarkKai He, Yucheng Huang, Wenqing Wang, Delong Ran et al.ACL 2025
- MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing AgentsYanqi Dai, Huanran Hu, Lei Wang, Shengjie Jin et al.ICLR 2025
- Speaker Verification in Agent-generated ConversationsYizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang et al.ACL 2024
- CoSER: Coordinating LLM-Based Persona Simulation of Established RolesXintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan et al.ICML 2025
