Evaluating Very Long-Term Conversational Memory of LLM Agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang
摘要
Existing works on long-term open-domain dialogues focus on evaluating model responses within contexts spanning no more than five chat sessions. Despite advancements in longcontext large language models (LLMs) and retrieval augmented generation (RAG) techniques, their efficacy in very long-term dialogues remains unexplored. To address this research gap, we introduce a machine-human pipeline to generate high-quality, very longterm dialogues by leveraging LLM-based agent architectures and grounding their dialogues on personas and temporal event graphs. Moreover, we equip each agent with the capability of sharing and reacting to images. The generated conversations are verified and edited by human annotators for long-range consistency and grounding to the event graphs. Using this pipeline, we collect LOCOMO, a dataset of very long-term conversations, each encompassing approx. 600 turns and 16K tokens on avg., over up to 32 sessions. Based on LOCOMO, we present a comprehensive evaluation benchmark to measure long-term memory in models, encompassing question answering, event summarization, and multi-modal dialogue generation tasks. Our experimental results indicate that LLMs exhibit challenges in understanding lengthy conversations and comprehending long-range temporal and causal dynamics within dialogues. Employing strategies like long-context LLMs or RAG can offer improvements but these models still substantially lag behind human performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper99
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao 等NeurIPS 2025 · 被引用 1,138 次
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 被引用 246 次
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemorySiru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen 等ICLR 2026 · 被引用 244 次
- LightMem: Lightweight and Efficient Memory-Augmented GenerationJizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang 等ICLR 2026 · 被引用 162 次
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie 等ACL 2026 · 被引用 140 次
它引用的顶会 Paper23
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
相关 Paper
- HyperMem: Hypergraph Memory for Long-Term ConversationsJuwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou 等ACL 2026 · 被引用 4 次
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 被引用 329 次
- ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional SupportTiantian Chen, Jiaqi Lu, Ying Shen, Lin ZhangWWW 2026 · 被引用 1 次
- SeCom: On Memory Construction and Retrieval for Personalized Conversational AgentsZhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo 等ICLR 2025
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMsMohammad Tavakoli, Alireza Salemi, Carrie Ye, Mohamed Abdalla 等ICLR 2026 · 被引用 56 次
