Benchmarking Data Science Agents
Yuge Zhang, Qiyang Jiang, Xingyu Han, Nan Chen, Yuqing Yang, Kan Ren
摘要
In the era of data-driven decision-making, the complexity of data analysis necessitates advanced expertise and tools of data science, presenting significant challenges even for specialists. Large Language Models (LLMs) have emerged as promising aids as data science agents, assisting humans in data analysis and processing. Yet their practical efficacy remains constrained by the varied demands of realworld applications and complicated analytical process. In this paper, we introduce DSEvala novel evaluation paradigm, as well as a series of innovative benchmarks tailored for assessing the performance of these agents throughout the entire data science lifecycle. Incorporating a novel bootstrapped annotation method, we streamline dataset preparation, improve the evaluation coverage, and expand benchmarking comprehensiveness. Our findings uncover prevalent obstacles and provide critical insights to inform future advancements in the field. * * Correspondence to Kan Ren. * Source code and data are available at https://github. com/MetaCopilot/dseval . Runtime Session In-memory Data country landArea pop2010 pop2023 pop2050 Calculate the population density for each country in 2023 and 2050. Result should be a new frame with "Country" as the index and "2023 Density" and "2050 Density" as the columns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren 等IEEE VIS 2024 · 被引用 44 次
- KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data LakesEugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Om Chabra 等ICLR 2026 · 被引用 37 次
- MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data ScienceRan Xu, Yuchen Zhuang, Yishan Zhong, Yue Yu 等ICLR 2026 · 被引用 17 次
- TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model CollaborationYuwei Du, Jie Feng, Jie Zhao, Yong LiNeurIPS 2025 · 被引用 8 次
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 被引用 3 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
相关 Paper
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao 等ICLR 2025
- Assessing and Verifying Task Utility in LLM-Powered ApplicationsNegar Arabzadeh, Siqing Huo, Nikhil Mehta, Qingyun Wu 等EMNLP 2024 · 被引用 7 次
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai 等ICML 2024 · 被引用 110 次
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang 等EMNLP 2024 · 被引用 7 次
- Lexara: A User-Centered Toolkit for Evaluating Large Language Models for Conversational Visual AnalyticsSrishti Palani, Vidya SetlurCHI 2026 · 被引用 1 次
