Subtle Bugs Everywhere: Generating Documentation for Data Wrangling Code
Chenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian Kästner
摘要
Data scientists reportedly spend a significant amount of their time in their daily routines on data wrangling, i.e. cleaning data and extracting features. However, data wrangling code is often repetitive and error-prone to write. Moreover, it is easy to introduce subtle bugs when reusing and adopting existing code, which results in reduced model quality. To support data scientists with data wrangling, we present a technique to generate documentation for data wrangling code. We use (1) program synthesis techniques to automatically summarize data transformations and (2) test case selection techniques to purposefully select representative examples from the data based on execution information collected with tailored dynamic program analysis. We demonstrate that a JupyterLab extension with our technique can provide on-demand documentation for many cells in popular notebooks and find in a user study that users with our plugin are faster and more effective at finding realistic bugs in data wrangling code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code VisualizationLiwenhan Xie, Chengbo Zheng, Haijun Xia, Huamin Qu 等UIST 2024 · 被引用 45 次
- Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and TraceabilityAvinash Bhat, Austin Coursey, Grace Hu, Sixian Li 等CHI 2023 · 被引用 29 次
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 被引用 25 次
- Revealing the Semantics of Data Wrangling Scripts With ComanticsKai Xiong, Zhongsu Luo, Siwei Fu, Yongheng Wang 等IEEE VIS 2022 · 被引用 13 次
- JupyterLab in Retrograde: Contextual Notifications That Highlight Fairness and Bias Issues for Data ScientistsGalen Harrison, Kevin Bryson, Ahmad Emmanuel Balla Bamba, Luca Dovichi 等CHI 2024 · 被引用 5 次
它引用的顶会 Paper7
- What's Wrong with Computational Notebooks? Pain Points, Needs, and Design OpportunitiesSouti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma 等CHI 2020 · 被引用 162 次
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine 等CHI 2020 · 被引用 110 次
- Assessing and Restoring Reproducibility of Jupyter NotebooksJiawei Wang, Tzu-yang Kuo, Li Li, Andreas ZellerASE 2020 · 被引用 68 次
- Restoring Execution Environments of Jupyter NotebooksJiawei Wang, Li Li, Andreas ZellerICSE 2021 · 被引用 49 次
- Callisto: Capturing the "Why" by Connecting Conversations with Computational NarrativesApril Yi Wang, Zihan Wu, Christopher Brooks, Steve OneyCHI 2020 · 被引用 44 次
相关 Paper
- Cell2Doc: ML Pipeline for Generating Documentation in Computational NotebooksTamal Mondal, Scott Barnett, Akash Lal, Jyothi VeduradaASE 2023 · 被引用 4 次
- Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling ScriptsYunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang 等CHI 2025 · 被引用 4 次
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee 等VLDB 2021 · 被引用 46 次
- Spine: Scaling up Programming-by-Negative-Example for String Filtering and TransformationChaoji Zuo, Sepehr Assadi, Dong DengSIGMOD 2022 · 被引用 4 次
- Bolt-on, Compact, and Rapid Program Slicing for Notebooks [Scalable Data Science]Shreya Shankar, Stephen Macke, Sarah E. Chasins, Andrew Head 等VLDB 2022 · 被引用 17 次
