REMem: Reasoning with Episodic Memory in Language Agent
Yiheng Shu, Padmaja Jonnalagedda, Xiang Gao, Bernal Jimenez Gutierrez, Weijian Qi, Kamalika Das, Huan Sun, Yu Su
Abstract
Humans excel at remembering concrete experiences along spatiotemporal contexts and performing reasoning across those events, i.e., the capacity for episodic memory. In contrast, memory in language agents remains mainly semantic, and current agents are not yet capable of effectively recollecting and reasoning over interaction histories. We identify and formalize the core challenges of episodic recollection and reasoning from this gap, and observe that existing work often overlooks episodicity, lacks explicit event modeling, or overemphasizes simple retrieval rather than complex reasoning. We present REMem, a two-phase framework for constructing and reasoning with episodic memory: 1) Offline indexing, where REMem converts experiences into a hybrid memory graph that flexibly links time-aware gists and facts. 2) Online inference, where REMem employs an agentic retriever with carefully curated tools for iterative retrieval over the memory graph. Comprehensive evaluation across four episodic memory benchmarks shows that REMem substantially outperforms state-of-the-art memory systems such as Mem0 and HippoRAG 2, showing 3.4% and 13.4% absolute improvements on episodic recollection and reasoning tasks, respectively. Moreover, REMem also demonstrates more robust refusal behavior for unanswerable questions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 246 citations
- Editing Large Language Models: Problems, Methods, and OpportunitiesYunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng et al.EMNLP 2023 · 83 citations
Related papers
- PlugMem: A Task-Agnostic Plugin Memory Module for LLM AgentsKe Yang, Zixi Chen, Xuan He, Jize Jiang et al.ICML 2026 · 20 citations
- ARTEM: Enhancing Large Language Model Agents with Spatial-Temporal Episodic MemoryCassandra Hui-Ming Tan, Budhitama Subagdja, Ah-Hwee TanAAAI 2026
- A Machine with Short-Term, Episodic, and Semantic Memory SystemsTaewoon Kim, Michael Cochez, Vincent François-Lavet, Mark A. Neerincx et al.AAAI 2023 · 8 citations
- HyperMem: Hypergraph Memory for Long-Term ConversationsJuwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou et al.ACL 2026 · 4 citations
- Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory UtilizationTaeyoon Kwon, Dongwook Choi, Hyojun Kim, Sunghwan Kim et al.ICLR 2026 · 14 citations
