CharacterBench: Benchmarking Character Customization of Large Language Models
Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Yuxuan Chen, Pei Ke, Zhuang Chen, Xiyao Xiao, Libiao Peng, Kuntian Tang, Rongsheng Zhang, Le Zhang
Abstract
Character-based dialogue (aka role-playing) enables users to freely customize characters for interaction, which often relies on LLMs, raising the need to evaluate LLMs' character customization capability. However, existing benchmarks fail to ensure a robust evaluation as they often only involve a single character category or evaluate limited dimensions. Moreover, the sparsity of character features in responses makes featurefocused generative evaluation both ineffective and inefficient. To address these issues, we propose CHARACTERBENCH, the largest bilingual generative benchmark, with 22,859 humanannotated samples covering 3,956 characters from 25 detailed character categories. We define 11 dimensions of 6 aspects, classified as sparse and dense dimensions based on whether character features evaluated by specific dimensions manifest in each response. We enable effective and efficient evaluation by crafting tailored queries for each dimension to induce characters' responses related to specific dimensions. Further, we develop CharacterJudge model for cost-effective and stable evaluations. Experiments show its superiority over SOTA automatic judges (e.g., GPT-4) and our benchmark's potential to optimize LLMs' character customization. Our repository is at https://github.com/thu-coai/CharacterBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 992fea94-b96f-4a13-8854-586c89d2facaCited by top-tier papers6
- Thinking in Character: Advancing Role-Playing Agents with Role-Aware ReasoningYihong Tang, Kehai Chen, Muyun Yang, Zheng-Yu Niu et al.NeurIPS 2025 · 16 citations
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 2 citations
- Persona-Pruner: Sculpting Lightweight Models for Role-PlayingJinsu Kim, Jihoon Tack, Noah Lee, Jongheon JeongICML 2026
- Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language ModelsZhouxing Tan, Ruochong Xiong, Yulong Wan, Jinlong Ma et al.AAAI 2026
- CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing AgentsYihong Tang, Kehai Chen, Liang Yue, Benyou Wang et al.ICML 2026
Builds on7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 97 citations
- ALOHA: Artificial Learning of Human Attributes for Dialogue AgentsAaron W. Li, Veronica Jiang, Steven Y. Feng, Julia Sprague et al.AAAI 2020 · 29 citations
- Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-AlignmentKeming Lu, Bowen Yu, Chang Zhou, Jingren ZhouACL 2024 · 16 citations
- InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan et al.ACL 2024
Related papers
- CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent EvaluationQuan Tu, Shilong Fan, Zihang Tian, Tianhao Shen et al.ACL 2024
- Crab: A Novel Configurable Role-Playing LLM with Assessing BenchmarkKai He, Yucheng Huang, Wenqing Wang, Delong Ran et al.ACL 2025
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li et al.AAAI 2025 · 6 citations
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackYouquan Li, Miao Zheng, Fan Yang, Guosheng Dong et al.EMNLP 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
