SWE-PDB: Teaching LLMs to Leverage Debugging Tools via Agentic Training
Jiaxing Liu, Xing Hu, Xin Xia
摘要
Interaction with debugging tools enables large language models (LLMs) to reason over concrete runtime states, rather than relying solely on static analysis of source code. Specifically, with the help of a debugger, an LLM-based agent can observe actual program execution by inspecting intermediate variable states and stepping through the control flow. These runtime observations enable the model to better understand program behavior and identify the root causes of bugs. Despite these advantages, using debuggers correctly and effectively remains challenging for many models because debugger interaction is inherently stateful and requires executing complex, long-horizon action sequences. As a result, models often exhibit unproductive interactions in which the debugger is underutilized or even disrupts the debugging process. To address this challenge, we propose SWE-PDB , the first training-based framework that teaches LLMs to leverage debuggers for interactive debugging and program repair. Our approach constructs large-scale buggy Python instances with verified failing tests from diverse sources and synthesizes multi-turn interactive debugging trajectories that follow structured debugging workflows. To ensure data quality, we apply multi-stage trajectory filtering and refinement, and train models using agentic supervised fine-tuning to learn effective debugger behaviors, followed by agentic reinforcement learning with rule-based rewards to improve generalization and promote more strategic debugger usage. Extensive evaluations across diverse benchmarks demonstrate substantial gains. In particular, SWE-PDB-14B achieves 38.0% accuracy on SWE-bench Verified with complete test suites, more than tripling the base model’s performance, while improving interaction efficiency and exhibiting robust, meaningful debugger usage.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- InspectCoder: Dynamic Analysis-Driven Self Repair through Interactive LLM-Debugger CollaborationYunkun Wang, Yue Zhang, Guochang Li, Chen Zhi 等OOPSLA 2026 · 被引用 1 次
- Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub ScenariosZhi Chen, Wei Ma, Lingxiao JiangICSE 2026
- LeDex: Training LLMs to Better Self-Debug and Explain CodeNan Jiang, Xiaopeng Li, Shiqi Wang, Qiang Zhou 等NeurIPS 2024 · 被引用 6 次
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han 等NeurIPS 2025 · 被引用 73 次
- Toward Training Superintelligent Software Agents through Self-Play SWE-RLYuxiang Wei, Zhiqing Sun, Emily McMilin, Jonas Gehring 等ICML 2026 · 被引用 32 次
