τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. Narasimhan
摘要
Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications. We propose τ -bench, a benchmark emulating dynamic conversations between a user (simulated by language models) and a language agent provided with domain-specific API tools and policy guidelines. We employ an efficient and faithful evaluation process that compares the database state at the end of a conversation with the annotated goal state. We also propose a new metric (pass^k) to evaluate the reliability of agent behavior over multiple trials. Our experiments show that even state-of-the-art function calling agents (like gpt-4o) succeed on < 50% of the tasks, and are quite inconsistent (pass^8 < 25% in retail). Our findings point to the need for methods that can improve the ability of agents to act consistently and follow rules reliably.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
相关 Paper
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li 等AAAI 2025 · 被引用 6 次
- GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM AgentsLingxiao Diao, Xinyue Xu, Wanxuan Sun, Cheng Yang 等ACL 2025
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based AgentsJiangyuan Wang, Kejun Xiao, Qi Sun, Huaipeng Zhao 等AAAI 2026 · 被引用 14 次
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan 等ICML 2026 · 被引用 15 次
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin 等EMNLP 2024 · 被引用 5 次
