Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
Ryan Louie, Ananjan Nandi, William Fang, Cheng Chang, Emma Brunskill, Diyi Yang
Abstract
Recent works leverage LLMs to roleplay realistic social scenarios, aiding novices in practicing their social skills. However, simulating sensitive interactions, such as in the domain of mental health, is challenging. Privacy concerns restrict data access, and collecting expert feedback, although vital, is laborious. To address this, we develop Roleplay-doh, a novel human-LLM collaboration pipeline that elicits qualitative feedback from a domain-expert, which is transformed into a set of principles, or natural language rules, that govern an LLM-prompted roleplay. We apply this pipeline to enable senior mental health supporters to create customized AI patients as simulated practice partners for novice counselors. After uncovering issues with basic GPT-4 simulations not adhering to expert-defined principles, we also introduce a novel principle-adherence prompting pipeline which shows a 30% improvement in response quality and principle following for the downstream task. Through a user study with 25 counseling experts, we demonstrate that the pipeline makes it easy and effective to create AI patients that more faithfully resemble real patients, as judged by both creators and third-party counselors. We provide access to the code and data on our project website: https://roleplay-doh.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2accfe2-516d-4282-a0e9-915ec4292518Cited by top-tier papers26
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari et al.CHI 2025 · 59 citations
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language ModelsLujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi et al.ICLR 2026 · 40 citations
- Scaffolding Empathy: Training Counselors with Simulated Patients and Utterance-level Performance VisualizationsIan Steenstra, Farnaz Nouraei, Timothy W. BickmoreCHI 2025 · 30 citations
- Cultural Learning-Based Culture Adaptation of Language ModelsChen Cecilia Liu, Anna Korhonen, Iryna GurevychACL 2025 · 14 citations
- Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice CounselorsRyan Louie, Raj Sanjay Shah, Ifdita Hasan Orney, Juan Pablo Pacheco et al.CHI 2026 · 7 citations
Builds on2
Related papers
- Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer CounselorsAlicja Chaszczewicz, Raj Sanjay Shah, Ryan Louie, Bruce A. Arnow et al.ACL 2024 · 12 citations
- 'Poker with Play Money': Exploring Psychotherapist Training with Virtual PatientsCynthia M. Baseman, Masum Hasan, Nathaniel Swinger, Sheila A. M. Rauch et al.CSCW 2025 · 2 citations
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained CounselorsZhiyang Qi, Takumasa Kaneko, Keiko Takamizo, Mariko Ukiyo et al.ACL 2025 · 7 citations
- PolicyPad: Collaborative Prototyping of LLM PoliciesK. J. Kevin Feng, Tzu-Sheng Kuo, Quan Ze Jim Chen, Inyoung Cheong et al.CHI 2026 · 2 citations
