Potemkin Understanding in Large Language Models
Marina Mancoridis, Bec Weeks, Keyon Vafa, Sendhil Mullainathan
Abstract
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated set of questions? This paper first introduces a formal framework to address this question. The key is to note that the benchmarks used to test LLMs-such as AP exams-are also those used to test people. However, this raises an implication: these benchmarks are only valid tests if LLMs misunderstand concepts in ways that mirror human misunderstandings. Otherwise, success on benchmarks only demonstrates potemkin understanding: the illusion of understanding driven by answers irreconcilable with how any human would interpret a concept. We present two procedures for quantifying the existence of potemkins: one using a specially designed benchmark in three domains, the other using a general procedure that provides a lower-bound on their prevalence. We find that potemkins are ubiquitous across models, tasks, and domains. We also find that these failures reflect not just incorrect understanding, but deeper internal incoherence in concept representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a35de20e-8879-4dcd-b67b-9837b7d15f61Cited by top-tier papers3
- The Serial Scaling HypothesisYuxi Liu, Konpat Preechakul, Kananart Kuwaranancharoen, Yutong BaiICLR 2026 · 12 citations
- Is In-Context Learning Learning?Adrian de WynterICLR 2026 · 1 citation
- The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to MistakesMarina Mancoridis, Zoe HitzigICML 2026
Builds on16
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
Related papers
- CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language ModelsHaibo Tong, Zeyang Yue, Feifei Zhao, Erliang Lin et al.ACL 2026
- Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language ModelsChani Jung, Dongkwan Kim, Jiho Jin, Jiseon Kim et al.EMNLP 2024 · 2 citations
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in InteractionsHyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras et al.EMNLP 2023 · 21 citations
- Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoningMelanie Sclar, Jane Dwivedi-Yu, Maryam Fazel-Zarandi, Yulia Tsvetkov et al.ICLR 2025
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 40 citations
