Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
Nitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek, Gal Yona
摘要
Standard factuality evaluations of LLMs treat all errors alike, obscuring whether failures arise from missing knowledge (empty shelves) or from limited access to encoded facts (lost keys). We propose a behavioral framework that profiles factual knowledge at the level of facts rather than questions, characterizing each fact by whether it is encoded, and then by how accessible it is: cannot be recalled, can be directly recalled, or can only be recalled with inference-time computation (thinking). To support such profiling, we introduce WikiProfile, a new benchmark constructed via an automated pipeline with a prompted LLM grounded in web search. Across 4 million responses from 13 LLMs, we find that encoding is nearly saturated in frontier models on our benchmark, with GPT-5 and Gemini-3 encoding 95-98% of facts. However, recall remains a major bottleneck: many errors previously attributed to missing knowledge instead stem from failures to access it. These failures are systematic and disproportionately affect long-tail facts and reverse questions. Finally, we show that thinking improves recall and can recover a substantial fraction of failures, indicating that future gains may rely less on scaling and more on methods that improve how models utilize what they already encode. Figure 1 | Top: We propose five knowledge profiles that characterize facts. Bottom: Percentages of these profiles across selected LLMs, revealing: (1) Scaling fills "empty shelves" by reducing encoding failures: frontier LLMs encode nearly all facts in our data. (2) Recall failures remain abundant despite scaling, leaving substantial room for improvement. (3) Thinking acts as a recovery mechanism of facts that would otherwise remain "lost".
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
相关 Paper
- OntoFact: Unveiling Fantastic Fact-Skeleton of LLMs via Ontology-Driven Reinforcement LearningZiyu Shang, Wenjun Ke, Nana Xiu, Peng Wang 等AAAI 2024 · 被引用 12 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMsYuxuan Li, Aoi Naito, Hirokazu ShiradoICML 2026 · 被引用 9 次
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen 等ICLR 2026 · 被引用 69 次
- FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationFarima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu WangACL 2025
