CoSER: Coordinating LLM-Based Persona Simulation of Established Roles
Xintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan, Rui Xu, Jen-tse Huang, Siyu Yuan, Haoran Guo, Jiangjie Chen, Shuchang Zhou, Wei Wang, Yanghua Xiao
Abstract
Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present COSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The COSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as conversation setups, character experiences and internal thoughts. Drawing from acting methodology, we introduce givencircumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop COSER 8B and COSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the COSER dataset for RPLA training, evaluation and retrieval. Moreover, COSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively. Our code, dataset and models are available at: https://github.com/Neph0s/CoSER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f85eef0-c0bd-48bb-988b-8df5fc433f6eCited by top-tier papers15
- Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-PlayingWenyuan Zhang, Shuaiyi Nie, Jiawei Sheng, Zefeng Zhang et al.EMNLP 2025 · 10 citations
- HumanLLM: Towards Personalized Understanding and Simulation of Human NatureYuxuan Lei, Tianfu Wang, Jianxun Lian, Zhengyu Hu et al.KDD 2026 · 4 citations
- Deriving Character Logic from Storyline as Codified Decision TreesLetian Peng, Kun Zhou, Longfei Yun, Yupeng Hou et al.ACL 2026 · 3 citations
- Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing AgentsYuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu et al.ICLR 2026 · 2 citations
- DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory GraphZhihao Xiao, Mengting Li, Xintao Wang, Linfeng Li et al.KDD 2026 · 1 citation
Builds on9
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 97 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-AlignmentKeming Lu, Bowen Yu, Chang Zhou, Jingren ZhouACL 2024 · 16 citations
Related papers
- CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based RewardsCheng Liu, Yifei Lu, Fanghua Ye, Jian Li et al.EMNLP 2025
- CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent EvaluationQuan Tu, Shilong Fan, Zihang Tian, Tianhao Shen et al.ACL 2024
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin et al.EMNLP 2024 · 2 citations
- CharacterBench: Benchmarking Character Customization of Large Language ModelsJinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi et al.AAAI 2025 · 7 citations
- Speaker Verification in Agent-generated ConversationsYizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang et al.ACL 2024
