DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Shaolei Zhang, Ju Fan, Meihao Fan, Yizhe Liu, Yuxin Zhang, Xiaoyong Du
Abstract
Autonomous data science, from raw data sources to analyst-grade deep research reports, has been a long-standing challenge, and is now becoming feasible with the emergence of powerful large language models (LLMs). Recent workflow-based data agents have shown promising results on specific data tasks but remain fundamentally limited in achieving fully autonomous data science due to their reliance on predefined workflows. In this paper, we introduce DeepAnalyze-8B, the first agentic LLM designed for autonomous data science, capable of automatically completing the end-toend pipeline from data sources to analyst-grade deep research reports. To tackle high-complexity data science tasks, we propose a curriculum-based agentic training paradigm that emulates the learning trajectory of human data scientists, enabling LLMs to progressively acquire and integrate multiple capabilities in real-world environments. We also introduce a data-grounded trajectory synthesis framework that constructs high-quality training data. Through agentic training, DeepAnalyze learns to perform a broad spectrum of data tasks, ranging from data question answering and specialized analytical tasks to open-ended data research. Experiments demonstrate that, with only 8B parameters, DeepAnalyze outperforms previous workflow-based agents built on most advanced proprietary LLMs. The model, code, and training data of DeepAnalyze are open-sourced, paving the way toward autonomous data science.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised RewardsYuxin Zhang, Meihao Fan, Ju Fan, Mingyang Yi et al.SIGMOD 2026 · 24 citations
- Scaling Generalist Data-Analytic AgentsShuofei Qiao, Yanqiu Zhao, Zhisong Qiu, Xiaobin Wang et al.ICLR 2026 · 12 citations
- DeepPrep: An LLM-Powered Agentic System for Autonomous Data PreparationMeihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang et al.VLDB 2026 · 4 citations
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context ManagementZherui Yang, Fan Liu, Yansong Ning, Hao LiuKDD 2026 · 3 citations
- CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang et al.ICML 2026 · 2 citations
Builds on10
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 210 citations
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 168 citations
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningSiyuan Guo, Cheng Deng, Ying Wen, Hechang Chen et al.ICML 2024 · 107 citations
Related papers
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsWei Liu, Peijie Yu, Michele Orini, Yali Du et al.ICML 2026 · 2 citations
- IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement LearningHaohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang et al.ICML 2026
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical StudyYuqi Zhu, Yi Zhong, Jintian Zhang, Ziheng Zhang et al.AAAI 2026 · 3 citations
- AndroidGen: Building an Android Language Agent under Data ScarcityHanyu Lai, Junjie Gao, Xiao Liu, Yifan Xu et al.ACL 2025
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
