TRACTUS: Understanding and Supporting Source Code Experimentation in Hypothesis-Driven Data Science
Krishna Subramanian, Johannes Maas, Jan O. Borchers
摘要
Data scientists experiment heavily with their code, compromising code quality to obtain insights faster. We observed ten data scientists perform hypothesis-driven data science tasks, and analyzed their coding, commenting, and analysis practice. We found that they have difficulty keeping track of their code experiments. When revisiting exploratory code to write production code later, they struggle to retrace their steps and capture the decisions made and insights obtained, and have to rerun code frequently. To address these issues, we designed TRACTUS, a system extending the popular RStudio IDE, that detects, tracks, and visualizes code experiments in hypothesis-driven data science tasks. TRACTUS helps recall decisions and insights by grouping code experiments into hypotheses, and structuring information like code execution output and documentation. Our user studies show how TRACTUS improves data scientists' workflows, and suggest additional opportunities for improvement. TRACTUS is available as an open source RStudio IDE addin at http://hci.rwth-aachen.de/tractus.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Critical Reflection on Visualization Research: Where Do Decision Making Tasks Hide?Evanthia Dimara, John T. StaskoIEEE VIS 2021 · 被引用 56 次
- Slide4N: Creating Presentation Slides from Computational Notebooks with Human-AI CollaborationFengjie Wang, Xuye Liu, Oujing Liu, Ali Neshati 等CHI 2023 · 被引用 37 次
- NBSearch: Semantic Search and Visual Exploration of Computational NotebooksXingjun Li, Yuanxin Wang, Hong Wang, Yang Wang 等CHI 2021 · 被引用 22 次
- OutlineSpark: Igniting AI-powered Presentation Slides Creation from Computational Notebooks through OutlinesFengjie Wang, Yanna Lin, Leni Yang, Haotian Li 等CHI 2024 · 被引用 19 次
- ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMsYuan Tian, Chuhan Zhang, Xiaotong Wang, Sitong Pan 等UIST 2025 · 被引用 3 次
相关 Paper
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 被引用 34 次
- Callisto: Capturing the "Why" by Connecting Conversations with Computational NarrativesApril Yi Wang, Zihan Wu, Christopher Brooks, Steve OneyCHI 2020 · 被引用 44 次
- On Understanding Data Worker Interaction BehaviorsLei Han, Tianwa Chen, Gianluca Demartini, Marta Indulska 等SIGIR 2020 · 被引用 12 次
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
- NoteFlow: Leveraging Charts as Sight Glasses for Consistent and Continuous Data Flow TracingYuan Tian, Dazhen Deng, Sen Yang, Huawei Zheng 等CHI 2026 · 被引用 1 次
