CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, Rui Yan
摘要
Recently, the advent of large language models 001 (LLMs) has revolutionized generative agents. 002 Among them, Role-Playing Conversational 003 Agents (RPCAs) attract considerable atten-004 tion due to their ability to emotionally engage 005 users. However, the absence of a compre-006 hensive benchmark impedes progress in this 007 field. To bridge this gap, we introduce Char-008 acterEval, a Chinese benchmark for compre-009 hensive RPCA assessment, complemented by a 010 tailored high-quality dataset. The dataset com-011 prises 1,785 multi-turn role-playing dialogues, 012 encompassing 11,376 examples and featuring 013 77 characters derived from Chinese novels and 014 scripts. It was carefully constructed, beginning 015 with initial dialogue extraction via GPT-4, fol-016 lowed by rigorous human-led quality control, 017 and enhanced with in-depth character profiles 018 sourced from Baidu Baike. CharacterEval em-019 ploys a multifaceted evaluation approach, en-020 compassing thirteen targeted metrics on four 021 dimensions. To facilitate the convenient eval-022 uation for these subjective metrics in Charac-023 terEval, we further developed CharacterRM, a 024 role-playing reward model based on human an-025 notations, which has a higher correlation with 026 human judgment compared to GPT-4. Compre-027 hensive experiments on CharacterEval demon-028 strate that Chinese LLMs exhibit more promis-029 ing capabilities than GPT-4 in Chinese role-030 playing conversation 1 . 031 1 Introduction 032 The development of large language models (LLMs) 033 has marked the beginning of a new era in conversa-034
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou 等ACL 2025 · 被引用 24 次
- Beyond Dialogue: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language ModelYeyong Yu, Runsheng Yu, Haojie Wei, Zhanqiu Zhang 等ACL 2025 · 被引用 13 次
- OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality InteractionHaonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu 等ACL 2025 · 被引用 10 次
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsMuyu He, Anand Kumar, Soumyadeep Bakshi, James Zou 等ACL 2026 · 被引用 8 次
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li 等AAAI 2025 · 被引用 6 次
它引用的顶会 Paper3
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue GenerationXiuyi Chen, Fandong Meng, Peng Li, Feilong Chen 等EMNLP 2020 · 被引用 78 次
- Zero-Resource Knowledge-Grounded Dialogue GenerationLinxiao Li, Can Xu, Wei Wu, Yufan Zhao 等NeurIPS 2020 · 被引用 75 次
相关 Paper
- CharacterBench: Benchmarking Character Customization of Large Language ModelsJinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi 等AAAI 2025 · 被引用 7 次
- Crab: A Novel Configurable Role-Playing LLM with Assessing BenchmarkKai He, Yucheng Huang, Wenqing Wang, Delong Ran 等ACL 2025
- MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing AgentsYanqi Dai, Huanran Hu, Lei Wang, Shengjie Jin 等ICLR 2025
- Speaker Verification in Agent-generated ConversationsYizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang 等ACL 2024
- CoSER: Coordinating LLM-Based Persona Simulation of Established RolesXintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan 等ICML 2025
