AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan, Zhongming Yu, Haozhou Xu, Lanxiang Hu, Abhilash Shankarampeta, Zimeng Huang, Wentao Ni, Yuandong Tian, Jishen Zhao
摘要
Large Language Models (LLMs) are increasingly used as autonomous agents in complex, long-horizon applications, where effective memory is critical for sustained performance. Yet existing memory benchmarks are largely dialogue-centric, while real agent memory consists of continuous agent-environment interaction trajectories composed of states, actions, observations, and tool outputs. To address this gap, we introduce AMA-Bench ( A gent M emory with A ny length), a benchmark for evaluating long-horizon memory in realistic agentic settings. AMA-Bench combines real-world agent trajectories from representative applications with expert-curated QA, as well as synthetic trajectories that scale to arbitrary horizons with rule-based QA. Our study shows that existing memory systems underperform because they fail to capture causal and objective information and rely heavily on lossy similarity-based retrieval. We further propose AMA-Agent , a memory system based on causality-graph construction and tool-augmented retrieval. AMA-Agent achieves 57.22% accuracy on AMA-Bench, outperforming the strongest baseline by 11.16% . Resources are available at: https://ama-bench.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao 等NeurIPS 2025 · 被引用 1,138 次
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 被引用 246 次
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon AgentsZijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim 等ICLR 2026 · 被引用 223 次
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM SystemsQingyao Ai, Yichen Tang, Changyue Wang, Jianming Long 等ICML 2026 · 被引用 47 次
相关 Paper
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu 等ICLR 2026 · 被引用 46 次
- Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous AgentsYiting Shen, Kun Li, Wei Zhou, Songlin HuACL 2026 · 被引用 12 次
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model AgentsYi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan 等ACL 2026 · 被引用 40 次
- UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon ScenariosHaotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang 等ICML 2026 · 被引用 21 次
- CloneMem: Benchmarking Long-Term Memory for AI ClonesSen Hu, Zhiyu Zhang, Yuxiang Wei, Xueran Han 等ACL 2026 · 被引用 6 次
