Code Code Evolution: Understanding How People Change Data Science Notebooks Over Time
Deepthi Raghunandan, Aayushi Roy, Shenzhi Shi, Niklas Elmqvist, Leilani Battle
摘要
Sensemaking is the iterative process of identifying, extracting, and explaining insights from data, where each iteration is referred to as the "sensemaking loop." Although recent work observes snapshots of the sensemaking loop within computational notebooks, none measure shifts in sensemaking behaviors over time-between exploration and explanation. This gap limits our ability to understand the full scope of the sensemaking process and thus our ability to design tools to fully support sensemaking. We contribute the first quantitative method to characterize how sensemaking evolves within data science computational notebooks. To this end, we conducted a quantitative study of 2,574 Jupyter notebooks mined from GitHub. First, we identify data science-focused notebooks that have undergone significant iterations. Second, we present regression models that automatically characterize sensemaking activity within individual notebooks by assigning them a score representing their position within the sensemaking spectrum. Finally, we use our regression models to calculate and analyze shifts in notebook scores across GitHub versions. Our results show that notebook authors participate in a diverse range of sensemaking tasks over time, such as annotation, branching analysis, and documentation. Finally, we propose design recommendations for extending notebook environments to support the sensemaking behaviors we observed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz StudyKen Gu, Madeleine Grunde-McLaughlin, Andrew M. McNutt, Jeffrey Heer 等CHI 2024 · 被引用 37 次
- Evaluating Navigation and Comparison Performance of Computational Notebooks on Desktop and in Virtual RealitySungwon In, Eric Krokos, Kirsten Whitley, Chris North 等CHI 2024 · 被引用 17 次
- Small, Medium, Large? A Meta-Study of Effect Sizes at CHI to Aid Interpretation of Effect Sizes and Power CalculationAnna-Marie Ortloff, Florin Martius, Mischa Meier, Theo Raimbault 等CHI 2025 · 被引用 17 次
- Enhancing Computational Notebooks with Code+Data Space VersioningHanxi Fang, Supawit Chockchowwat, Hari Sundaram, Yongjoo ParkCHI 2025 · 被引用 6 次
- Pagebreaks: Multi-Cell Scopes in Computational NotebooksEric Rawn, Sarah E. ChasinsCHI 2025 · 被引用 4 次
它引用的顶会 Paper3
- What's Wrong with Computational Notebooks? Pain Points, Needs, and Design OpportunitiesSouti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma 等CHI 2020 · 被引用 162 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
- Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data AnalysisYang Liu, Tim Althoff, Jeffrey HeerCHI 2020 · 被引用 40 次
相关 Paper
- Loops: Leveraging Provenance and Visualization to Support Exploratory Data Analysis in NotebooksKlaus Eckelt, Kiran Gadhave, Alexander Lex, Marc StreitIEEE VIS 2024 · 被引用 7 次
- NBSearch: Semantic Search and Visual Exploration of Computational NotebooksXingjun Li, Yuanxin Wang, Hong Wang, Yang Wang 等CHI 2021 · 被引用 22 次
- Characterizing Practices, Limitations, and Opportunities Related to Text Information Extraction Workflows: A Human-in-the-loop PerspectiveSajjadur Rahman, Eser KandoganCHI 2022 · 被引用 20 次
- How Scientists Use Jupyter Notebooks: Goals, Quality Attributes, and OpportunitiesRuanqianqian (Lisa) Huang, Savitha Ravi, Michael He, Boyu Tian 等ICSE 2025 · 被引用 3 次
- NoteFlow: Leveraging Charts as Sight Glasses for Consistent and Continuous Data Flow TracingYuan Tian, Dazhen Deng, Sen Yang, Huawei Zheng 等CHI 2026 · 被引用 1 次
