Improving data scientist efficiency with provenance
Jingmei Hu, Jiwon Joung, Maia L. Jacobs, Krzysztof Z. Gajos, Margo I. Seltzer
Abstract
Data scientists frequently analyze data by writing scripts. We conducted a contextual inquiry with interdisciplinary researchers, which revealed that parameter tuning is a highly iterative process and that debugging is time-consuming. As analysis scripts evolve and become more complex, analysts have difficulty conceptualizing their workflow. In particular, after editing a script, it becomes difficult to determine precisely which code blocks depend on the edit. Consequently, scientists frequently re-run entire scripts instead of re-running only the necessary parts. We present ProvBuild, a tool that leverages language-level provenance to streamline the debugging process by reducing programmer cognitive load and decreasing subsequent runtimes, leading to an overall reduction in elapsed debugging time. ProvBuild uses provenance to track dependencies in a script. When an analyst debugs a script, ProvBuild generates a simplified script that contains only the information necessary to debug a particular problem. We demonstrate that debugging the simplified script lowers a programmer's cognitive load and permits faster re-execution when testing changes. The combination of reduced cognitive load and shorter runtime reduces the time necessary to debug a script. We quantitatively and qualitatively show that even though ProvBuild introduces overhead during a script's first execution, it is a more efficient way for users to debug and tune complex workflows. ProvBuild demonstrates a novel use of language-level provenance, in which it is used to proactively improve programmer productively rather than merely providing a way to retroactively gain insight into a body of code. CCS CONCEPTS • Software and its engineering → Software development techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 25 citations
- Demystifying "bad" error messages in data science librariesYida Tao, Zhihui Chen, Yepang Liu, Jifeng Xuan et al.FSE 2021 · 6 citations
Related papers
- Enabling Efficient Attack Investigation via Human-in-the-Loop Security AnalysisSaimon Amanuel Tsegai, Xinyu Yang, Haoyuan Liu, Peng GaoVLDB 2025 · 2 citations
- Online and Interactive Bayesian Inference DebuggingNathanael Nussbaumer, Markus Böck, Jürgen CitoICSE 2026
- BugDoc: Algorithms to Debug Computational ProcessesRaoni Lourenço, Juliana Freire, Dennis E. ShashaSIGMOD 2020 · 9 citations
- Everything Everyway All at Once - Time Traveling Debugging for Stream Processing ApplicationsTimo Räth, Marius Schlegel, Kai-Uwe SattlerICDE 2024 · 1 citation
- Modus: a Datalog dialect for building container imagesChris Tomy, Tingmao Wang, Earl T. Barr, Sergey MechtaevFSE 2022 · 3 citations
