AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Junzhi Chen, Harsh Trivedi, Jane Pan, Michael Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal
Abstract
Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark's difficulty and its potential to advance research on user-in-the-loop tool-use agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li et al.ICLR 2026 · 520 citations
- TaleBrush: Sketching Stories with Generative Pretrained Language ModelsJohn Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee et al.CHI 2022 · 202 citations
- Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMsHariharan Subramonyam, Roy Pea, Christopher Lawrence Pondoc, Maneesh Agrawala et al.CHI 2024 · 137 citations
- Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily AssistantGaole He, Gianluca Demartini, Ujwal GadirajuCHI 2025 · 91 citations
Related papers
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world ApplicationsWei He, Yueqing Sun, Hongyan Hao, Xueyuan Hao et al.ICLR 2026 · 41 citations
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Material StructuresTaoyuze Lv, Alexander Chen, Fengyu Xie, Chu Wu et al.ICML 2026 · 3 citations
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsQuyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao et al.ACL 2026 · 38 citations
