Scaling Generalist Data-Analytic Agents
Shuofei Qiao, Yanqiu Zhao, Zhisong Qiu, Xiaobin Wang, Jintian Zhang, Zhao Bin, Ningyu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen
Abstract
Data-analytic agents are emerging as a key catalyst for automated scientific discovery and for the vision of Innovating AI. Current approaches, however, rely heavily on prompt engineering or multi-agent scaffolds over proprietary models, while open-source models still struggle with diverse-format, large-scale data files and long-horizon, multi-step reasoning that real-world analytics demands. This paper introduces DATAMIND, a scalable data synthesis and agent training recipe designed to construct generalist data-analytic agents. DATAMIND tackles three key challenges in building open-source data-analytic agents, including insufficient data resources, improper training strategy, and unstable code-based multi-turn rollout. Concretely, DATAMIND applies 1) a fine-grained task taxonomy and a recursive easy-to-hard task composition mechanism to increase the diversity and difficulty of synthesized queries; 2) a knowledge-augmented trajectory sampling strategy followed by model-based and rule-based filtering; 3) a dynamically adjustable training objective combining both SFT and RL losses; 4) a memory-frugal and stable code-based multi-turn rollout framework. Built on DATAMIND, we curate DATAMIND-12K, a high-quality trajectory set spanning diverse domains, task categories, and data file formats for data-analytic tasks. Trained on DATAMIND-12K, our DATAMIND-14B achieves state-of-the-art with an average score of 71.16% on multiple data analysis benchmarks, outperforming the strongest proprietary baselines DeepSeek-V3.1 and GPT-5. Our DATAMIND-7B also performs best among all open-source models with a score of 68.10%. We also list some empirical insights gained from our exploratory trials in the analysis experiments, aiming to provide actionable insights about agent training for the community. We have released DATAMIND-12K and DATAMIND-7B,14B for the community 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5253586-5997-41b9-a599-24b8da5f4a7fCited by top-tier papers5
- DSGym: A Standardized and Holistic Framework for Evaluating and Training Data Science AgentsFan Nie, Junlin Wang, Harper Hua, Federico Bianchi et al.ICML 2026 · 12 citations
- Can We Predict Before Executing Machine Learning Agents?Jingsheng Zheng, Jintian Zhang, Yujie Luo, Yuren Mao et al.ACL 2026 · 6 citations
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context ManagementZherui Yang, Fan Liu, Yansong Ning, Hao LiuKDD 2026 · 3 citations
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsWei Liu, Peijie Yu, Michele Orini, Yali Du et al.ICML 2026 · 2 citations
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning ProblemShuofei Qiao, Yunxiang Wei, Xuehai Wang, Bin Wu et al.ICML 2026
Builds on31
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
Related papers
- DeepAnalyze: Agentic Large Language Models for Autonomous Data ScienceShaolei Zhang, Ju Fan, Meihao Fan, Yizhe Liu et al.ICML 2026 · 48 citations
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang et al.EMNLP 2024 · 7 citations
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao et al.ICLR 2025
- DAComp: Benchmarking Data Agents across the Full Data Intelligence LifecycleFangyu Lei, Jinxiang Meng, Yiming Huang, Junjie zhao et al.ICLR 2026 · 18 citations
