Subtle Bugs Everywhere: Generating Documentation for Data Wrangling Code
Chenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian Kästner
Abstract
Data scientists reportedly spend a significant amount of their time in their daily routines on data wrangling, i.e. cleaning data and extracting features. However, data wrangling code is often repetitive and error-prone to write. Moreover, it is easy to introduce subtle bugs when reusing and adopting existing code, which results in reduced model quality. To support data scientists with data wrangling, we present a technique to generate documentation for data wrangling code. We use (1) program synthesis techniques to automatically summarize data transformations and (2) test case selection techniques to purposefully select representative examples from the data based on execution information collected with tailored dynamic program analysis. We demonstrate that a JupyterLab extension with our technique can provide on-demand documentation for many cells in popular notebooks and find in a user study that users with our plugin are faster and more effective at finding realistic bugs in data wrangling code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1f527ca-eed7-45bf-8ebf-e1ea4ed87ff2Cited by top-tier papers8
- WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code VisualizationLiwenhan Xie, Chengbo Zheng, Haijun Xia, Huamin Qu et al.UIST 2024 · 45 citations
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li et al.CHI 2023 · 29 citations
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 25 citations
- Revealing the Semantics of Data Wrangling Scripts With ComanticsKai Xiong, Zhongsu Luo, Siwei Fu, Yongheng Wang et al.IEEE VIS 2022 · 13 citations
- JupyterLab in Retrograde: Contextual Notifications That Highlight Fairness and Bias Issues for Data ScientistsGalen Harrison, Kevin Bryson, Ahmad Emmanuel Balla Bamba, Luca Dovichi et al.CHI 2024 · 5 citations
Builds on7
- What's Wrong with Computational Notebooks? Pain Points, Needs, and Design OpportunitiesSouti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma et al.CHI 2020 · 162 citations
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine et al.CHI 2020 · 110 citations
- Assessing and Restoring Reproducibility of Jupyter NotebooksJiawei Wang, Tzu-yang Kuo, Li Li, Andreas ZellerASE 2020 · 68 citations
- Restoring Execution Environments of Jupyter NotebooksJiawei Wang, Li Li, Andreas ZellerICSE 2021 · 49 citations
- Callisto: Capturing the "Why" by Connecting Conversations with Computational NarrativesApril Yi Wang, Zihan Wu, Christopher Brooks, Steve OneyCHI 2020 · 44 citations
Related papers
- Cell2Doc: ML Pipeline for Generating Documentation in Computational NotebooksTamal Mondal, Scott Barnett, Akash Lal, Jyothi VeduradaASE 2023 · 4 citations
- Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling ScriptsYunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang et al.CHI 2025 · 4 citations
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee et al.VLDB 2021 · 46 citations
- Spine: Scaling up Programming-by-Negative-Example for String Filtering and TransformationChaoji Zuo, Sepehr Assadi, Dong DengSIGMOD 2022 · 4 citations
- Bolt-on, Compact, and Rapid Program Slicing for Notebooks [Scalable Data Science]Shreya Shankar, Stephen Macke, Sarah E. Chasins, Andrew Head et al.VLDB 2022 · 17 citations
