Simulacrum of Stories: Examining Large Language Models as Qualitative Research Participants
Shivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari, Sarah E. Fox
Abstract
The recent excitement around generative models has sparked a wave of proposals suggesting the replacement of human participation and labor in research and development–e.g., through surveys, experiments, and interviews—with synthetic research data generated by large language models (LLMs). We conducted interviews with 19 qualitative researchers to understand their perspectives on this paradigm shift. Initially skeptical, researchers were surprised to see similar narratives emerge in the LLM-generated data when using the interview probe. However, over several conversational turns, they went on to identify fundamental limitations, such as how LLMs foreclose participants’ consent and agency, produce responses lacking in palpability and contextual depth, and risk delegitimizing qualitative research methods. We argue that the use of LLMs as proxies for participants enacts the surrogate effect, raising ethical and epistemological concerns that extend beyond the technical limitations of current models to the core of whether LLMs fit within qualitative ways of knowing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4d9848e-f3b8-42e5-a7d8-599b94b6d39bCited by top-tier papers9
- Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public OpinionsJoseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang et al.ACL 2025 · 48 citations
- A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language TechnologiesAlicia DeVrio, Myra Cheng, Lisa Egede, Alexandra Olteanu et al.CHI 2025 · 28 citations
- Spectrum Tuning: Post-Training for Distributional Coverage and In-Context SteerabilityTaylor Sorensen, Benjamin Newman, Jared Moore, Chan Young Park et al.ICLR 2026 · 19 citations
- Towards Better Reflexive Thematic Analysis in HCI: A Scoping Review of Practice at CHIJacob M. Rigby, Ioanna IacovidesCHI 2026 · 6 citations
- Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 ArticlesDanial Amin, Joni Salminen, Farhan Ahmed, Sonja M. H. Tervola et al.CHI 2026 · 6 citations
Builds on28
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
Related papers
- Large Language Models in Qualitative Research: Uses, Tensions, and IntentionsHope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno et al.CHI 2025 · 40 citations
- Interview-Informed Generative Agents for Product Discovery: A Validation StudyZichao Wang, Alexa F. SiuCHI 2026 · 1 citation
- Collecting Qualitative Data at Scale with Large Language Models: A Case StudyAlejandro Cuevas Villalba, Jennifer V. Scurrell, Eva Maxfield Brown, Jason Entenmann et al.CSCW 2025 · 14 citations
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li et al.CSCW 2025 · 31 citations
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 6 citations
