Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
Sanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap, Graham Neubig
摘要
AI agents are increasingly being deployed to automate tasks, often based on underspecified user instructions. Making unwarranted assumptions to compensate for the missing information and failing to ask clarifying questions can lead to suboptimal outcomes, safety risks due to tool misuse, and wasted computational resources. In this work, we study the ability of LLM agents to handle underspecified instructions in interactive code generation settings by evaluating proprietary and open-weight models on their performance across three key steps: (a) detecting underspecificity, (b) asking targeted clarification questions, and (c) leveraging the interaction to improve performance in underspecified scenarios. We introduce Ambig-SWE, an underspecified variant of SWE-Bench Verified, specifically designed to evaluate agent behavior under ambiguity and interaction. Our findings reveal that models struggle to distinguish between well-specified and underspecified instructions. However, when models interact for underspecified inputs, they effectively obtain vital information from the user leading to significant improvements in performance, up to 74% over the non-interactive settings, underscoring the value of effective interaction. Our study highlights critical gaps in how current state-of-the-art models handle missing information in complex software engineering tasks and structures the evaluation into distinct steps to enable targeted improvements 1 . 1 Code and data can be accessed at https://github.com/sani903/InteractiveSWEAgents
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TOM-SWE: User Mental Modeling For Software Engineering AgentsXuhui Zhou, Valerie Chen, Zhiruo Wang, Graham Neubig 等ICML 2026 · 被引用 12 次
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsJaylen Jones, Zhehao Zhang, Yuting Ning, Eric Fosler-Lussier 等ICML 2026 · 被引用 9 次
- Morae: Proactively Pausing UI Agents for User ChoicesYi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy PavelUIST 2025 · 被引用 4 次
- Code with Me or for Me? How Increasing AI Automation Transforms Developer WorkflowsValerie Chen, Ameet Talwalkar, Robert Brennan, Graham NeubigCHI 2026 · 被引用 2 次
- How can we assess human-agent interactions? Case studies in software agent designValerie Chen, Rohit Malhotra, Xingyao Wang, Juan Michelini 等ICML 2026
它引用的顶会 Paper9
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang 等ICLR 2024 · 被引用 288 次
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- Goal Driven Discovery of Distributional Differences via Language DescriptionsRuiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn 等NeurIPS 2023 · 被引用 81 次
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu 等ICLR 2025 · 被引用 7 次
相关 Paper
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Uncertainty-Aware Clarification in LLM Agents with Information GainMengyi DENG, Zhiwei Li, Xin Li, Tingyu ZHU 等ICML 2026 · 被引用 1 次
- Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering TasksDimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski 等AAAI 2026 · 被引用 1 次
- Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the WildAayush Kumar, Yasharth Bajpai, Sumit Gulwani, Gustavo Soares 等ASE 2025 · 被引用 1 次
