Virtual Personas for Language Models via an Anthology of Backstories
Suhong Moon, Marwa Abdulhai, Minwoo Kang, Joseph Suh, Widyadewi Soedarmadji, Eran Kohen Behar, David M. Chan
摘要
Large language models (LLMs) are trained from vast repositories of text authored by millions of distinct authors, reflecting an enormous diversity of human traits. While these models bear the potential to be used as approximations of human subjects in behavioral studies, prior efforts have been limited in steering model responses to match individual human users. In this work, we introduce "Anthology", a method for conditioning LLMs to particular virtual personas by harnessing open-ended life narratives, which we refer to as "backstories." We show that our methodology enhances the consistency and reliability of experimental outcomes while ensuring better representation of diverse subpopulations. Across three nationally representative human surveys conducted as part of Pew Research Center's American Trends Panel (ATP), we demonstrate that Anthology achieves up to 18% improvement in matching the response distributions of human respondents and 27% improvement in consistency metrics. Our code is available at https://github.com/CannyLab/anthology . A: I'm 37. I grew up in a small town, in a small house … A: Certainly! I am a new college grad from New Jersey … A: Born and raised in Tennessee, I had a blissful childhood … Step 1. LLM-Generation of Backstories Step 3. Match Virtual Personas to Target Human User Distribution Match to Human User Distribution (Demographic Variables) Step 2. Demographic Survey on Virtual Personas Q: What is your age? (a) 18-29 (d) 65 or above (b) 30-49 (d) Prefer not to answer (c) 50-64 A: (b) 37 years old. Backstory Conditioned Virtual Persona Q: What is the highest level of education you have completed ? (a) Less than high school … … A: (e) Bachelor's degree LLM Q: Tell me about yourself.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff 等NeurIPS 2025 · 被引用 51 次
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public OpinionsJoseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang 等ACL 2025 · 被引用 48 次
- Enhancing Personalized Multi-Turn Dialogue with Curiosity RewardYanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani 等NeurIPS 2025 · 被引用 32 次
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou 等ACL 2025 · 被引用 24 次
- Finetuning LLMs for Human Behavior Prediction in Social Science ExperimentsAkaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. BernsteinEMNLP 2025 · 被引用 12 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- Social Simulacra: Creating Populated Prototypes for Social Computing SystemsJoon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris 等UIST 2022 · 被引用 192 次
相关 Paper
- Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsGeorg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus StrohmaierACL 2026 · 被引用 4 次
- Human Simulacra: Benchmarking the Personification of Large Language ModelsQiujie Xie, Qiming Feng, Tianqi Zhang, Qingqiu Li 等ICLR 2025
- HumanLM: Simulating Users with State Alignment Beats Response ImitationShirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang 等ICML 2026
- Synthia: Scalable Grounded Persona Generation from Social Media DataVahid Rahimzadeh, Erfan Moosavi Monazzah, Mohammad Taher Pilehvar, Yadollah YaghoobzadehACL 2026 · 被引用 1 次
- Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona DescriptionsJoni Salminen, Chang Liu, Wenjing Pian, Jianxing Chi 等CHI 2024 · 被引用 55 次
