Flipping the Dialogue: Training and Evaluating User Language Models
Tarek Naous, Philippe Laban, Wei Xu, Jennifer Neville
Abstract
Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user's request. To satisfy this specific role, LMs are post-trained to be helpful assistants -optimized to produce exhaustive and well-structured responses, free of ambiguity and grammar errors. User utterances, on the other hand, are rarely perfected, with each user phrasing requests in unique ways, sometimes putting in partial effort at each turn and refining on the fly. To evaluate LM performance in realistic settings, prior work simulated users in multi-turn conversations, often by prompting an LM originally trained to be a helpful assistant to act as a user. However, we show that assistant LMs make for poor user simulators, with the surprising finding that better assistants yield worse simulators. Instead, we introduce purpose-built User Language Models (User LMs) -models post-trained to simulate human users in multi-turn conversations. Through various evaluations, we show how User LMs align better with human behavior and achieve better simulation robustness than existing simulation methods. When leveraging User LMs to simulate coding and math conversations, the performance of a strong assistant (GPT-4o) drops from 74.6% to 57.4%, confirming that more realistic simulation environments lead to assistant struggles as they fail to cope with the nuances of users in multi-turn setups. microsoft/UserLM-8b SIMULATING USERS IN CONVERSATIONS … … USING AN ASSISTANT LANGUAGE MODEL User Intent: Write a Python function: given an array of integers, sort ones between 1 and 9 inclusive, reverse the array, and replace digits by their name from "One", "Two", "Three", etc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f363fe7-fb95-4b85-83c4-e681d43a67c9Cited by top-tier papers5
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang et al.ICML 2026 · 3 citations
- Thinking Alignment of Scenario-Oriented User SimulationXiaoting Wu, Yi Huang, Chunyang Gao, Mengfei Guo et al.ACL 2026
- Persona-Pruner: Sculpting Lightweight Models for Role-PlayingJinsu Kim, Jihoon Tack, Noah Lee, Jongheon JeongICML 2026
- HumanLM: Simulating Users with State Alignment Beats Response ImitationShirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang et al.ICML 2026
- DiscoverLLM: From Executing Intents to Discovering ThemTae Soo Kim, Yoonjoo Lee, Jaesang Yu, John Chung et al.ICML 2026
Builds on14
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public OpinionsJoseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang et al.ACL 2025 · 48 citations
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
Related papers
- SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie et al.EMNLP 2025
- PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User SimulatorChuyi Kong, Yaxin Fan, Xiang Wan, Feng Jiang et al.ACL 2024 · 3 citations
- An Empirical Study of Python Library Migration Using Large Language ModelsMohayeminul Islam, Ajay Kumar Jha, May Mahmoud, Ildar Akhmetov et al.ASE 2025 · 2 citations
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun et al.EMNLP 2024 · 10 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
