Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
Dimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski, Robert West, Dimitrios Gunopulos
Abstract
Standard single-turn, static benchmarks fall short in evaluating the nuanced capabilities of Large Language Models (LLMs) on complex tasks such as software engineering. In this work, we propose a novel interactive evaluation framework that assesses LLMs on multi-requirement programming tasks through structured, feedback-driven dialogue. Each task is modeled as a requirement dependency graph, and an "interviewer" LLM, aware of the ground-truth solution, provides minimal, targeted hints to an "interviewee" model to help correct errors and fulfill target constraints. This dynamic protocol enables fine-grained diagnostic insights into model behavior, uncovering strengths and systematic weaknesses that static benchmarks fail to measure. We build on DevAI, a benchmark of 55 curated programming tasks, by adding ground-truth solutions and evaluating the relevance and utility of interviewer hints through expert annotation. Our results highlight the importance of dynamic evaluation in advancing the development of collaborative code-generating agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a396d6a3-bbf2-4f5f-b469-7c1454f6b2d9Builds on6
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong et al.ICLR 2024 · 92 citations
- IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringRuosen Li, Ruochen Li, Barry Wang, Xinya DuNeurIPS 2024 · 26 citations
Related papers
- INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and HelpfulnessHung Le, Doyen Sahoo, Yingbo Zhou, Caiming Xiong et al.NeurIPS 2024 · 12 citations
- Talk2Code: A Multi-Turn Interaction Benchmark with Dual-Track Evaluation for Code GenerationWeibin Yang, Liangru Xie, Jieyun Cai, Yuxiang Yan et al.AAAI 2026
- Agent-as-a-Judge: Evaluate Agents with AgentsMingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang et al.ICML 2025
- Enhancing Code Generation via Bidirectional Comment-Level Mutual GroundingYifeng Di, Tianyi ZhangICSE 2025 · 3 citations
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen et al.ICLR 2026 · 25 citations
