Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors
Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
Abstract
LLMs are transforming software development, yet most code benchmarks still emphasize syntactic or functional correctness in simple, single-error cases. These settings miss the core difficulty of real-world data science debugging, where runtime errors propagate across multiple lines (multi-hop) and often appear in sets (multi-bug). We introduce DSDBench: Data Science Debugging Benchmark, the first benchmark to systematically evaluate LLMs on this challenge. Unlike general debugging benchmark suites such as SWE-bench, DSD-Bench targets non-expert, data-centric scripting, where practitioners rely heavily on blackbox libraries and write exploratory code that is error-prone and difficult to debug. Evaluations of state-of-the-art LLMs reveal large performance gaps: even frontier models that excel at code generation fail to reliably trace and resolve these errors, exposing a critical "generation versus understanding" gap. DSDBench provides a resource to drive progress toward more robust and trustworthy AI-assisted data science. 1 cause_error_line: y_pred = model.predict(X_train) effect_error_line (different from cause): mse = mean_squared_error(y_test, y_pred) error_message: ValueError: Found input variables with inconsistent numbers of samples cause_error_line: X = imputer.fit_transform(y) effect_error_line (different from cause): model.fit(X_train, y_train) error_message: ValueError: Input y contains NaN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e27bee7e-1971-43a8-b455-4a1e13667a8bBuilds on4
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng et al.ICML 2024 · 73 citations
- On the self-verification limitations of large language models on reasoning and planning tasksKaya Stechly, Karthik Valmeekam, Subbarao KambhampatiICLR 2025
Related papers
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun et al.AAAI 2026 · 10 citations
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao et al.ICLR 2025
- Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub ScenariosZhi Chen, Wei Ma, Lingxiao JiangICSE 2026
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang et al.EMNLP 2024 · 7 citations
- Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?Jing Ye, Yiwen Duan, Yonghong Yu, Victor Ma et al.ICML 2026 · 1 citation
