Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, Estefany Kelly Buchanan, Junhong Shen, Guanghao Ye
Abstract
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 965e3cee-06b1-4fa4-960c-2117762d4b71Cited by top-tier papers14
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan et al.ICML 2026 · 15 citations
- RExBench: Can coding agents autonomously implement AI research extensions?Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin et al.ACL 2026 · 10 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
- EvoClaw: Evaluating AI Agents on Continuous Software EvolutionGangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan et al.ICML 2026 · 6 citations
- VeRO: A Harness for Agents to Optimize AgentsVarun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue et al.ICML 2026 · 6 citations
Builds on14
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng et al.ICLR 2026 · 27 citations
Related papers
- BioAgent Bench: An AI Agent Evaluation Suite for BioinformaticsDionizije Fa, Marko Culjak, Bruno Pandza, Mateo CupicICML 2026 · 7 citations
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding AgentsJingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang et al.ICML 2026 · 37 citations
- FrontierCS: Evolving Challenges for Evolving IntelligenceQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao et al.ICML 2026
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsChristopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz et al.ICLR 2025
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
