-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan
摘要
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce -bench, with four key contributions: 1. A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2. A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3. A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4. fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, -bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir 等ICLR 2026 · 被引用 86 次
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGymWeihua Du, Hailei Gong, Zhan Ling, Kang Liu 等ICLR 2026 · 被引用 13 次
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 被引用 5 次
- Scaling Agentic Capabilities via Grounded Interaction SynthesisWenhang Shi, Jinhao Dong, Yiren Chen, Zhe Zhao 等KDD 2026 · 被引用 1 次
- Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented DialogueNing Gao, Wei Zhang, Yuqin Dai, Ling Shi 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper4
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis 等ICLR 2024 · 被引用 292 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
相关 Paper
- Non-Collaborative User Simulators for Tool AgentsJeonghoon Shim, Woojung Song, Cheyon Jin, Seungwon KooK 等ICLR 2026 · 被引用 17 次
- -Voice: Benchmarking Full-Duplex Voice Agents on Real-World DomainsSoham Ray, Keshav Dhandhania, Victor Barres, Karthik NarasimhanICML 2026
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan 等ICML 2026 · 被引用 15 次
- TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic ReasoningSina Tayebati, Divake Kumar, Nastaran Darabi, Davide Ettori 等ICML 2026 · 被引用 7 次
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
