-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan
Abstract
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce -bench, with four key contributions: 1. A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2. A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3. A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4. fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, -bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 102e620c-13ea-4050-9f1a-57e7e939ce64Cited by top-tier papers9
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGymWeihua Du, Hailei Gong, Zhan Ling, Kang Liu et al.ICLR 2026 · 13 citations
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 5 citations
- Scaling Agentic Capabilities via Grounded Interaction SynthesisWenhang Shi, Jinhao Dong, Yiren Chen, Zhe Zhao et al.KDD 2026 · 1 citation
- Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented DialogueNing Gao, Wei Zhang, Yuqin Dai, Ling Shi et al.ICML 2026 · 1 citation
Builds on4
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Identifying the Risks of LM Agents with an LM-Emulated SandboxYangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis et al.ICLR 2024 · 292 citations
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
Related papers
- Non-Collaborative User Simulators for Tool AgentsJeonghoon Shim, Woojung Song, Cheyon Jin, Seungwon KooK et al.ICLR 2026 · 17 citations
- -Voice: Benchmarking Full-Duplex Voice Agents on Real-World DomainsSoham Ray, Keshav Dhandhania, Victor Barres, Karthik NarasimhanICML 2026
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan et al.ICML 2026 · 15 citations
- TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic ReasoningSina Tayebati, Divake Kumar, Nastaran Darabi, Davide Ettori et al.ICML 2026 · 7 citations
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
