AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
Keyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li
Abstract
Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture long-horizon real-world scenarios. Moreover, the reliance on human-inthe-loop feedback for realistic tasks creates a scalability bottleneck, hindering automated rollout collection and evaluation. To bridge this gap, we introduce AGENCYBENCH, a comprehensive benchmark derived from daily AI usage, evaluating 6 core agentic capabilities across 32 real-world scenarios, comprising 138 tasks with specific queries, deliverables, and rubrics. These scenarios require an average of 90 tool calls, 1 million tokens, and hours of execution time to resolve. To enable automated evaluation, we employ a user simulation agent to provide iterative feedback, and a Docker sandbox to conduct visual and functional rubric-based assessment. Experiments reveal that closed-source models significantly outperform opensource models (48.4% vs 32.1%). Further analysis reveals significant disparities across models in resource efficiency, feedback-driven self-correction, and specific tool-use preferences. Finally, we investigate the impact of agentic scaffolds, observing that proprietary models demonstrate superior performance within their native ecosystems (e.g., Claude-4.5-Opus via Claude-Agent-SDK), while open-source models exhibit distinct performance peaks, suggesting potential optimization for specific execution frameworks. AGENCYBENCH serves as a critical testbed for next-generation agents, highlighting the necessity of co-optimizing model architecture with agentic frameworks. We believe this work sheds light on the future direction of autonomous agents, and to facilitate community adoption, we release the full benchmark and evaluation toolkit at https://github.com/GAIR-NLP/AgencyBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5abaec4-3843-4489-a866-e5b4b3d999bdCited by top-tier papers3
- daVinci-Dev: Agent-native Mid-training for Software EngineeringJi Zeng, Dayuan Fu, Tiantian Mi, Zhuang Yumin et al.ICML 2026 · 13 citations
- Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksZexue He, Yu Wang, Churan Zhi, Yuanzhe Hu et al.ICML 2026 · 2 citations
- QiMeng-LibBench: Benchmarking LLM Agents for Library-Scale Cross-Architecture MigrationWeijia Li, KE GAO, Jiajie Li, Han Sun et al.ICML 2026
Builds on7
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim et al.ICLR 2026 · 154 citations
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP UseZijian Wu, Xiangyan Liu, Xinyuan Zhang, Lingjun Chen et al.ICLR 2026 · 40 citations
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song et al.ACL 2025 · 12 citations
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu et al.ICLR 2025 · 7 citations
- SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time ScalingYang Xiao, Chunpu Xu, Ruifeng Yuan, Jessie Wang et al.AAAI 2026 · 1 citation
Related papers
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao et al.ICLR 2026 · 30 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- ShortcutsBench: A Large-Scale Real-world Benchmark for API-based AgentsHaiyang Shen, Yue Li, Desong Meng, Dongqi Cai et al.ICLR 2025
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai et al.ACL 2025
- SoMe: A Realistic Benchmark for LLM-based Social Media AgentsDizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu et al.AAAI 2026 · 1 citation
