Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function
Keyon Vafa, Ashesh Rambachan, Sendhil Mullainathan
摘要
What makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they will be used for. We consider a setting where these deployment decisions are made by people, and in particular, people's beliefs about where an LLM will perform well. We model such beliefs as the consequence of a human generalization function: having seen what an LLM gets right or wrong, people generalize to where else it might succeed. We collect a dataset of 19K examples of how humans make generalizations across 79 tasks from the MMLU and BIG-Bench benchmarks. We show that the human generalization function can be predicted using NLP methods: people have consistent structured ways to generalize. We then evaluate LLM alignment with the human generalization function. Our results show that -- especially for cases where the cost of mistakes is high -- more capable models (e.g. GPT-4) can do worse on the instances people choose to use them for, exactly because they are not aligned with the human generalization function.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language ModelsJessica Y. Bo, Sophia Wan, Ashton AndersonCHI 2025 · 被引用 31 次
- A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice ExperimentsManuel Cherep, Chengtian Ma, Abigail Xu, Maya Shaked 等ICLR 2026 · 被引用 13 次
- What's Producible May Not Be Reachable: Measuring the Steerability of Generative ModelsKeyon Vafa, Sarah Bentley, Jon M. Kleinberg, Sendhil MullainathanNeurIPS 2025 · 被引用 5 次
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMsZongjie Li, Daoyuan Wu, Shuai Wang, Zhendong SuCCS 2025 · 被引用 1 次
- People Can Accurately Predict Behavior of Complex Algorithms That Are Available, Compact, and Aligned CSCW031Lindsay Popowski, Helena Vasconcelos, Ignacio Javier Fernandez, Chijioke Chinaza Mgbahurike 等CSCW 2026
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 被引用 962 次
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
相关 Paper
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 被引用 11 次
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma 等FSE 2024 · 被引用 8 次
- Large Language Models Assume People are More Rational than We Really areRyan Liu, Jiayi Geng, Joshua C. Peterson, Ilia Sucholutsky 等ICLR 2025
- On Evaluating LLM Alignment by Evaluating LLMs as JudgesYixin Liu, Pengfei Liu, Arman CohanNeurIPS 2025 · 被引用 7 次
- Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansJen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren 等NeurIPS 2024 · 被引用 63 次
