DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python
Jinglin Peng, Weiyuan Wu, Brandon Lockhart, Song Bian, Jing Nathan Yan, Linghao Xu, Zhixuan Chi, Jeffrey M. Rzeszotarski, Jiannan Wang
摘要
Exploratory Data Analysis (EDA) is a crucial step in any data science project. However, existing Python libraries fall short in supporting data scientists to complete common EDA tasks for statistical modeling. Their API design is either too low level, which is optimized for plotting rather than EDA, or too high level, which is hard to specify more fine-grained EDA tasks. In response, we propose DataPrep.EDA, a novel task-centric EDA system in Python. Dat-aPrep.EDA allows data scientists to declaratively specify a wide range of EDA tasks in different granularity with a single function call. We identify a number of challenges to implement Dat-aPrep.EDA, and propose effective solutions to improve the scalability, usability, customizability of the system. In particular, we discuss some lessons learned from using Dask to build the data processing pipelines for EDA tasks and describe our approaches to accelerate the pipelines. We conduct extensive experiments to compare DataPrep.EDA with Pandas-profiling, the state-of-the-art EDA system in Python. The experiments show that DataPrep.EDA significantly outperforms Pandas-profiling in terms of both speed and user experience. DataPrep.EDA is open-sourced as an EDA component of DataPrep: https://github.com/sfu-db/dataprep.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dead or Alive: Continuous Data Profiling for Interactive Data ScienceWill Epperson, Vaishnavi Gorantla, Dominik Moritz, Adam PererIEEE VIS 2023 · 被引用 28 次
- Can Large Language Models Predict Data Correlations from Column Names?Immanuel TrummerVLDB 2023 · 被引用 17 次
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao 等VLDB 2025 · 被引用 10 次
- KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceMossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier 等ICDE 2024 · 被引用 7 次
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesQixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui 等SIGMOD 2025 · 被引用 4 次
它引用的顶会 Paper2
相关 Paper
- Dias: Dynamic Rewriting of Pandas CodeStefanos Baziotis, Daniel D. Kang, Charith MendisSIGMOD 2024 · 被引用 7 次
- Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science EngineWeizheng Lu, Chao Hui, Yunhai Wang, Feng Zhang 等VLDB 2025 · 被引用 1 次
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark 等VLDB 2022 · 被引用 61 次
- Reliable and Cost-Effective Exploratory Data Analysis via Graph-Guided RAGMossad Helali, Yutai Luo, Tae Jun Ham, Jim Plotts 等EMNLP 2025
- PipelineProfiler: A Visual Analytics Tool for the Exploration of AutoML PipelinesJorge Piazentin Ono, Sonia Castelo, Roque Lopez, Enrico Bertini 等IEEE VIS 2020 · 被引用 49 次
