Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors
Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
摘要
LLMs are transforming software development, yet most code benchmarks still emphasize syntactic or functional correctness in simple, single-error cases. These settings miss the core difficulty of real-world data science debugging, where runtime errors propagate across multiple lines (multi-hop) and often appear in sets (multi-bug). We introduce DSDBench: Data Science Debugging Benchmark, the first benchmark to systematically evaluate LLMs on this challenge. Unlike general debugging benchmark suites such as SWE-bench, DSD-Bench targets non-expert, data-centric scripting, where practitioners rely heavily on blackbox libraries and write exploratory code that is error-prone and difficult to debug. Evaluations of state-of-the-art LLMs reveal large performance gaps: even frontier models that excel at code generation fail to reliably trace and resolve these errors, exposing a critical "generation versus understanding" gap. DSDBench provides a resource to drive progress toward more robust and trustworthy AI-assisted data science. 1 cause_error_line: y_pred = model.predict(X_train) effect_error_line (different from cause): mse = mean_squared_error(y_test, y_pred) error_message: ValueError: Found input variables with inconsistent numbers of samples cause_error_line: X = imputer.fit_transform(y) effect_error_line (different from cause): model.fit(X_train, y_train) error_message: ValueError: Input y contains NaN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama 等ICML 2024 · 被引用 270 次
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai 等ICML 2024 · 被引用 110 次
- NExT: Teaching Large Language Models to Reason about Code ExecutionAnsong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng 等ICML 2024 · 被引用 73 次
- On the self-verification limitations of large language models on reasoning and planning tasksKaya Stechly, Karthik Valmeekam, Subbarao KambhampatiICLR 2025
相关 Paper
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun 等AAAI 2026 · 被引用 10 次
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao 等ICLR 2025
- Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub ScenariosZhi Chen, Wei Ma, Lingxiao JiangICSE 2026
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang 等EMNLP 2024 · 被引用 7 次
- Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?Jing Ye, Yiwen Duan, Yonghong Yu, Victor Ma 等ICML 2026 · 被引用 1 次
