Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Wu, Siru Ouyang, Tom Tang, Jiaxin Pei, Julian McAuley
Abstract
Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversations or text but fails to capture how memory is used to guide future decisions. Another class focuses on agent acting in single-session tasks without the need for long-term memory. However, in realistic settings, memorization and action are tightly coupled: agents acquire memory while interacting with the environment, and subsequently rely on that memory to solve future tasks. To capture this setting, We introduce MEMORYARENA, a unified evaluation gym for benchmarking agent memory in multi-session Memory-Agent-Environment loops. The benchmark consists of human-crafted agentic tasks with explicitly interdependent subtasks, where agents must learn from earlier actions and feedback by distilling experiences into memory, and subsequently use that memory to guide later actions to solve the overall task. MEMORYARENA supports evaluation across web navigation, preference-constrained planning, progressive information searching, and sequential formal reasoning, and reveals that agents with near-saturated performance on existing longcontext memory benchmarks like LoCoMo perform poorly in our agentic setting, exposing a gap in current evaluations for agents with memory. MEMORYARENA is released at https: //memoryarena.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f793cc00-5f17-47d9-8092-4a6bc01e6b1eCited by top-tier papers1
Ask how each one uses itBuilds on7
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 246 citations
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemorySiru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen et al.ICLR 2026 · 244 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM SystemsQingyao Ai, Yichen Tang, Changyue Wang, Jianming Long et al.ICML 2026 · 47 citations
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal et al.ACL 2024 · 30 citations
Related papers
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
- Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM AgentsYuanchen Bei, Tianxin Wei, Xuying Ning, Yanjun Zhao et al.ACL 2026 · 23 citations
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement LearningEgor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. PanovICLR 2026 · 43 citations
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji et al.ICML 2024 · 188 citations
- Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM AgentsYifei Li, Weidong Guo, Lingling Zhang, Rongman Xu et al.ACL 2026 · 5 citations
