Do Large Language Models Know How Much They Know?
Gabriele Prato, Jerry Huang, Prasanna Parthasarathi, Shagun Sodhani, Sarath Chandar
Abstract
Large Language Models (LLMs) have emerged as highly capable systems and are increasingly being integrated into various uses.Nevertheless, the rapid advancement in their deployment trails a comprehensive understanding of their internal mechanisms, as well as a delineation of their capabilities and limitations.A desired characteristic of an intelligent system is its ability to recognize the scope of its own knowledge.To investigate whether LLMs embody this attribute, we develop a benchmark that challenges these models to enumerate all information they possess on specific topics.This benchmark assesses whether the models recall excessive, insufficient, or the precise amount of required information, thereby indicating their awareness of how much they know about the given topic.Our findings reveal that the emergence of this property varies across different architectures and manifests at diverse rates.However, with sufficient scaling, all tested models are ultimately capable of performing this task.The insights gained from this research advance our understanding of LLMs, shedding light on their operational capabilities and contributing to the ongoing exploration of their intricate dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 263d41dd-63b6-45a4-9c83-198c86f113feBuilds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
Related papers
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li et al.ACL 2025 · 33 citations
- KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsYuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan et al.WWW 2024 · 6 citations
- ALCUNA: Large Language Models Meet New KnowledgeXunjian Yin, Baizhou Huang, Xiaojun WanEMNLP 2023 · 5 citations
- Working Memory Identifies Reasoning Limits in Language ModelsChunhui Zhang, Yiren Jian, Zhongyu Ouyang, Soroush VosoughiEMNLP 2024 · 4 citations
- Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge ExtractionYuheng Yang, Siqi Zhu, Tao Feng, Ge Liu et al.ICML 2026
