EpiK-Eval: Evaluation for Language Models as Epistemic Models
Gabriele Prato, Jerry Huang, Prasanna Parthasarathi, Shagun Sodhani, Sarath Chandar
摘要
In the age of artificial intelligence, the role of large language models (LLMs) is becoming increasingly central. Despite their growing prevalence, their capacity to consolidate knowledge from different training documents-a crucial ability in numerous applications-remains unexplored. This paper presents the first study examining the capability of LLMs to effectively combine such information within their parameter space. We introduce EpiK-Eval, a novel question-answering benchmark tailored to evaluate LLMs' proficiency in formulating a coherent and consistent knowledge representation from segmented narratives. Evaluations across various LLMs reveal significant weaknesses in this domain. We contend that these shortcomings stem from the intrinsic nature of prevailing training objectives. Consequently, we advocate for refining the approach towards knowledge consolidation, as it harbors the potential to dramatically improve their overall effectiveness and performance. The findings from this study offer insights for developing more robust and reliable LLMs. Our code and benchmark are available at https: //github.com/chandar-lab/EpiK-Eval
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
相关 Paper
- Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM ReasoningZiqing Zhuang, Linhai Zhang, Jiasheng Si, Deyu Zhou 等ACL 2026
- Episodic Memories Generation and Evaluation Benchmark for Large Language ModelsAlexis Huet, Zied Ben-Houidi, Dario RossiICLR 2025
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin 等EMNLP 2023 · 被引用 21 次
- KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsYuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan 等WWW 2024 · 被引用 6 次
- PhantomWiki: On-Demand Datasets for Reasoning and Retrieval EvaluationAlbert Gong, Kamile Stankeviciute, Chao Wan, Anmol Kabra 等ICML 2025
