UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon Scenarios
Haotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang, Zeyu Qin, Wenjie Lu, Guozheng Ma, Haiying He, Yingsha Xie, Qiyang Zhou, Zixuan Hu, Hongze Mi
Abstract
Autonomous agents have recently achieved remarkable progress across diverse domains, yet most evaluations focus on short-horizon, fully observable tasks. In contrast, many critical real-world tasks, such as large-scale software development, commercial investment, and scientific discovery, unfold in long-horizon and partially observable scenarios where success hinges on sustained reasoning, planning, memory management, and tool use. Existing benchmarks rarely capture these long-horizon challenges, leaving a gap in systematic evaluation. To bridge this gap, we introduce , a novel benchmark that measures the foundational capabilities essential for complex real-world challenges. We use exploration as a unifying task across three distinct environments to validate these core competencies. Agents are designed in long-horizon discovery tasks where they must iteratively uncover hidden rules through sustained reasoning, planning, memory and tools management, and interaction with environments. Under the heaviest scale setting, trajectories average tokens and tool calls, whereas in standard configurations they still exceed tokens and involve more than tool calls on average. Our extensive experiments reveal that agents powered by state-of-the-art LLMs consistently underperform in these settings, whereas human participants achieve much higher scores, underscoring a persistent gap in agents' long-horizon exploration abilities. We also observe that simple scaling fails in our task. To better illustrate the failure of agents, we conduct an in-depth analysis of collected trajectories. We identify eight types of errors and attribute them to two primary causes: in-context locking and functional fundamental capability gaps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 294de106-dce4-44f4-a15c-8ccc38cc6e23Builds on6
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han et al.NeurIPS 2025 · 73 citations
- AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental LearningMinghao Chen, Yihang Li, Yanting Yang, Shiyu Yu et al.NeurIPS 2024 · 67 citations
Related papers
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu et al.ACL 2026 · 21 citations
- AMA-Bench: Evaluating Long-Horizon Memory for Agentic ApplicationsYujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan et al.ICML 2026 · 40 citations
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On GamesDavide Paglieri, Bartlomiej Cupial, Samuel Coward, Ulyana Piterbarg et al.ICLR 2025
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding AgentsJingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang et al.ICML 2026 · 37 citations
- ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool SandboxYuanyang Li, Xue Yang, Longyue Wang, Weihua Luo et al.ICML 2026 · 1 citation
