Episodic Memories Generation and Evaluation Benchmark for Large Language Models
Alexis Huet, Zied Ben-Houidi, Dario Rossi
摘要
Episodic memory -the ability to recall specific events grounded in time and space -is a cornerstone of human cognition, enabling not only coherent storytelling, but also planning and decision-making. Despite their remarkable capabilities, Large Language Models (LLMs) lack a robust mechanism for episodic memory: we argue that integrating episodic memory capabilities into LLM is essential for advancing AI towards human-like cognition, increasing their potential to reason consistently and ground their output in real-world episodic events, hence avoiding confabulations. To address this challenge, we introduce a comprehensive framework to model and evaluate LLM episodic memory capabilities. Drawing inspiration from cognitive science, we develop a structured approach to represent episodic events, encapsulating temporal and spatial contexts, involved entities, and detailed descriptions. We synthesize a unique episodic memory benchmark, free from contamination, and release open source code and datasets to assess LLM performance across various recall and episodic reasoning tasks. Our evaluation of state-of-the-art models, including GPT-4 and Claude variants, Llama 3.1, and o1-mini, reveals that even the most advanced LLMs struggle with episodic memory tasks, particularly when dealing with multiple related events or complex spatio-temporal relationships -even in contexts as short as 10k-100k tokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory UtilizationTaeyoon Kwon, Dongwook Choi, Hyojun Kim, Sunghwan Kim 等ICLR 2026 · 被引用 14 次
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic WorkspacesShreyas Rajesh, Pavan Holur, Chenda Duan, David Chong 等AAAI 2026 · 被引用 3 次
- Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User HistorySerin Kim, Sangam Lee, Dongha LeeICML 2026
- (How) Do Language Models Track State?Belinda Z. Li, Zifan Carl Guo, Jacob AndreasICML 2025
- ARTEM: Enhancing Large Language Model Agents with Spatial-Temporal Episodic MemoryCassandra Hui-Ming Tan, Budhitama Subagdja, Ah-Hwee TanAAAI 2026
它引用的顶会 Paper14
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
相关 Paper
- Human-inspired Episodic Memory for Infinite Context LLMsZafeirios Fountas, Martin Benfeghoul, Adnan Oomerjee, Fenia Christopoulou 等ICLR 2025 · 被引用 1 次
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao 等NeurIPS 2025 · 被引用 30 次
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEvalIda Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma 等NeurIPS 2023 · 被引用 114 次
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu 等ACL 2024 · 被引用 12 次
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
