Uncertainty Quantification for LLM-Based Survey Simulations
Chengpiao Huang, Yuhang Wu, Kaizheng Wang
摘要
We investigate the use of large language models (LLMs) to simulate human responses to survey questions, and perform uncertainty quantification to assess the fidelity of the simulations. Our approach converts imperfect black-box LLMsimulated responses into confidence sets for population parameters of human responses. A key innovation lies in determining the optimal number of simulated responses: too many produce overly narrow confidence sets with poor coverage, while too few yield excessively loose estimates. Our method adaptively selects the simulation sample size that ensures valid average-case coverage guarantees. The selected sample size itself further provides a quantitative measure of LLM-human misalignment. Experiments on real survey datasets reveal heterogeneous fidelity gaps across different LLMs and domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 被引用 116 次
- Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading EfficiencyEric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang 等EMNLP 2023 · 被引用 4 次
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsNitay Calderon, Roi Reichart, Rotem DrorACL 2025
相关 Paper
- Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language ModelsGeorg Ahnert, Anna-Carolina Haensch, Barbara Plank, Markus StrohmaierACL 2026 · 被引用 4 次
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn 等ICML 2026 · 被引用 24 次
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and RectificationStefan Krsteski, Giuseppe Russo, Serina Chang, Robert West 等ACL 2026 · 被引用 10 次
- Estimating Semantic Alphabet Size for LLM Uncertainty QuantificationLucas H. McCabe, Rimon Melamed, Tom Hartvigsen, H. Howie HuangICLR 2026 · 被引用 7 次
- Prediction-Powered Ranking of Large Language ModelsIvi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 被引用 34 次
