Lune

NeurIPS2025Top-tier venue

Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning

Marwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff, Sergey Levine, Natasha Jaques

2025Year
51Citations
4Top-tier citations

Abstract

Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics-prompt-to-line consistency, line-to-line consistency, and Q&A consistency-that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multiturn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent, faithful, and trustworthy simulated users.

2 Related Work

The promise of LLMs as proxies for human behavior encourages their adoption as scalable simulations of social interaction for use in fields such as psychology, education, political science, and AI alignment [71,50]. These models are used not merely as impersonal chatbots, but as stand-ins for students, patients, voters, and citizens. Their behaviors can shape downstream AI applications or guide the training of agentic systems. LLMs are well-suited to in this role due to their fluency, generality, and responsiveness to conditioning, but ensuring that these simulated agents are realistic remains a major open challenge. Treating LLMs as human simulators requires not only mastering world modeling-the ability to predict and generate contextually appropriate language-but also

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 05c34005-2646-49e7-ae1e-192ff2dc415c

Cited by top-tier papers4

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines