ACL2026

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models

Wenlong Shi, Jianxun Lian, Mingqi Wu, Haiming Qin, Mingyang Zhou, Xing Xie, Naipeng Chao, Hao Liao

摘要

Large language models (LLMs) increasingly serve as interactive social agents, yet their ability to maintain coherent and authentic personalevel role-playing remains limited, particularly in realistic social scenarios. Existing research predominantly focuses on characterlevel settings and relies on static evaluation formats, failing to capture the complexity of everyday social interactions. In this work, we present PersonaArena, a dynamic simulation framework for evaluating and improving persona-level role-playing in LLMs. Person-aArena leverages a large, filtered corpus of usergenerated social content to construct a nuanced persona bank, and elicits multi-turn, contextrich interactions within simulated social environments. Our framework features a multiagent debating judge for holistic and unbiased assessment. Through extensive experiments, we demonstrate that PersonaArena enables rigorous evaluation and enhancement of LLMs' role-playing capabilities, advancing the development of more authentic and socially adept AI agents. The code for the PersonaArena framework is available at our public GitHub repository: https://aka.ms/personaarena .