SWE-PDB: Teaching LLMs to Leverage Debugging Tools via Agentic Training
Jiaxing Liu, Xing Hu, Xin Xia
Abstract
Interaction with debugging tools enables large language models (LLMs) to reason over concrete runtime states, rather than relying solely on static analysis of source code. Specifically, with the help of a debugger, an LLM-based agent can observe actual program execution by inspecting intermediate variable states and stepping through the control flow. These runtime observations enable the model to better understand program behavior and identify the root causes of bugs. Despite these advantages, using debuggers correctly and effectively remains challenging for many models because debugger interaction is inherently stateful and requires executing complex, long-horizon action sequences. As a result, models often exhibit unproductive interactions in which the debugger is underutilized or even disrupts the debugging process. To address this challenge, we propose SWE-PDB , the first training-based framework that teaches LLMs to leverage debuggers for interactive debugging and program repair. Our approach constructs large-scale buggy Python instances with verified failing tests from diverse sources and synthesizes multi-turn interactive debugging trajectories that follow structured debugging workflows. To ensure data quality, we apply multi-stage trajectory filtering and refinement, and train models using agentic supervised fine-tuning to learn effective debugger behaviors, followed by agentic reinforcement learning with rule-based rewards to improve generalization and promote more strategic debugger usage. Extensive evaluations across diverse benchmarks demonstrate substantial gains. In particular, SWE-PDB-14B achieves 38.0% accuracy on SWE-bench Verified with complete test suites, more than tripling the base model’s performance, while improving interaction efficiency and exhibiting robust, meaningful debugger usage.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- InspectCoder: Dynamic Analysis-Driven Self Repair through Interactive LLM-Debugger CollaborationYunkun Wang, Yue Zhang, Guochang Li, Chen Zhi et al.OOPSLA 2026 · 1 citation
- Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub ScenariosZhi Chen, Wei Ma, Lingxiao JiangICSE 2026
- LeDex: Training LLMs to Better Self-Debug and Explain CodeNan Jiang, Xiaopeng Li, Shiqi Wang, Qiang Zhou et al.NeurIPS 2024 · 6 citations
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han et al.NeurIPS 2025 · 73 citations
- Toward Training Superintelligent Software Agents through Self-Play SWE-RLYuxiang Wei, Zhiqing Sun, Emily McMilin, Jonas Gehring et al.ICML 2026 · 32 citations
