DSCodeBench: A Realistic Benchmark for Data Science Code Generation
Shuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun, Qihao Zhu, Jie M. Zhang
Abstract
We introduce DSCodeBench, a new benchmark designed to evaluate large language models (LLMs) on complicated and realistic data science code generation tasks. DSCodeBench consists of 1,000 carefully constructed problems sourced from realistic problems from GitHub across ten widely used Python data science libraries. DSCodeBench offers a more challenging and representative testbed, more complex code solutions, more comprehensive data science libraries, clearer and better structured problem descriptions, and stronger test suites. To construct the DSCodeBench, we develop a robust pipeline that combines task scope selection, code construction, test case generation, and problem description synthesis. The process is paired with rigorous manual editing to ensure alignment and enhance the reliability of the evaluation. Experimental result shows that DSCodeBench exhibits robust scaling behavior, where larger models systematically outperform smaller ones, validating its ability to distinguish model capabilities. The best LLM we test, GPT-4o, has a pass@1 of 0.392, indicating that LLMs still have a large room to improve for realistic data science code generation tasks. We believe DSCodeBench will serve as a rigorous and trustworthy foundation for advancing LLM-based data science programming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang et al.ICML 2026 · 2 citations
- Bridging Instead of Replacing Online Coding Communities with AI through Community-Enriched Chatbot Designs CSCW008Junling Wang, Lahari Goswami, Gustavo Kreia Umbelino, Kiara Chau et al.CSCW 2026
Builds on14
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionWei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang et al.NeurIPS 2024 · 210 citations
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 71 citations
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang et al.FSE 2024 · 68 citations
- EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationDong Huang, Jianbo Dai, Han Weng, Puzhen Wu et al.NeurIPS 2024 · 54 citations
Related papers
- RealisticCodeBench: Towards More Realistic Evaluation of Large Language Models for Code GenerationXiao Yu, Haoxuan Chen, Lei Liu, Xing Hu et al.ASE 2025
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao et al.ICLR 2025
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang et al.EMNLP 2024 · 7 citations
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkDewu Zheng, Yanlin Wang, Ensheng Shi, Xilin Liu et al.ICSE 2026
- AutoCodeBench: Large Language Models are Automatic Code Benchmark GeneratorsChangzhi Zhou, Ao Liu, Yuchi Deng, Zhiying Zeng et al.ICLR 2026 · 27 citations
