The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-Large
Sumon Biswas, Mohammad Wardat, Hridesh Rajan
摘要
Increasingly larger number of software systems today are including data science components for descriptive, predictive, and prescriptive analytics. The collection of data science stages from acquisition, to cleaning/curation, to modeling, and so on are referred to as data science pipelines. To facilitate research and practice on data science pipelines, it is essential to understand their nature. What are the typical stages of a data science pipeline? How are they connected? Do the pipelines differ in the theoretical representations and that in the practice? Today we do not fully understand these architectural characteristics of data science pipelines. In this work, we present a three-pronged comprehensive study to answer this for the stateof-the-art, data science in-the-small, and data science in-the-large. Our study analyzes three datasets: a collection of 71 proposals for data science pipelines and related concepts in theory, a collection of over 105 implementations of curated data science pipelines from Kaggle competitions to understand data science in-the-small, and a collection of 21 mature data science projects from GitHub to understand data science in-the-large. Our study has led to three representations of data science pipelines that capture the essence of our subjects in theory, in-the-small, and in-the-large. CCS CONCEPTS • Software and its engineering → Software creation and management; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Towards Understanding Fairness and its Composition in Ensemble Machine LearningUsman Gohar, Sumon Biswas, Hridesh RajanICSE 2023 · 被引用 30 次
- Dango: A Mixed-Initiative Data Wrangling System using Large Language ModelWei-Hao Chen, Weixi Tong, Amanda Case, Tianyi ZhangCHI 2025 · 被引用 19 次
- Design by Contract for Deep Learning APIsShibbir Ahmed, Sayem Mohammad Imtiaz, Syeda Khairunnesa Samantha, Breno Dantas Cruz 等FSE 2023 · 被引用 10 次
- Reflecting on Design Paradigms of Animated Data Video ToolsLeixian Shen, Haotian Li, Yun Wang, Huamin QuCHI 2025 · 被引用 10 次
- Enabling Secure and Efficient Data Analytics Pipeline Evolution with Trusted Execution EnvironmentHaotian Gao, Cong Yue, Tien Tuan Anh Dinh, Zhiyong Huang 等VLDB 2023 · 被引用 6 次
它引用的顶会 Paper8
- Repairing deep neural networks: fix patterns and challengesMd Johirul Islam, Rangeet Pan, Giang Nguyen, Hridesh RajanICSE 2020 · 被引用 102 次
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 被引用 101 次
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 被引用 96 次
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 被引用 93 次
- Restoring Execution Environments of Jupyter NotebooksJiawei Wang, Li Li, Andreas ZellerICSE 2021 · 被引用 49 次
相关 Paper
- NB2P: Generating Data Science Pipelines from Computational NotebooksHaotian Gao, Quang Trung Ta, Tien Tuan Anh Dinh, Nhut-Minh Ho 等ICSE 2026
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- A Large-Scale Study of Model Integration in ML-Enabled Software SystemsYorick Sens, Henriette Knopp, Sven Peldszus, Thorsten BergerICSE 2025 · 被引用 3 次
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 等ICSE 2022 · 被引用 11 次
- The Product Beyond the Model - An Empirical Study of Repositories of Open-Source ML ProductsNadia Nahar, Haoran Zhang, Grace A. Lewis, Shurui Zhou 等ICSE 2025 · 被引用 2 次
