Dead or Alive: Continuous Data Profiling for Interactive Data Science
Will Epperson, Vaishnavi Gorantla, Dominik Moritz, Adam Perer
Abstract
Profiling data by plotting distributions and analyzing summary statistics is a critical step throughout data analysis. Currently, this process is manual and tedious since analysts must write extra code to examine their data after every transformation. This inefficiency may lead to data scientists profiling their data infrequently, rather than after each transformation, making it easy for them to miss important errors or insights. We propose continuous data profiling as a process that allows analysts to immediately see interactive visual summaries of their data throughout their data analysis to facilitate fast and thorough analysis. Our system, AutoProfiler, presents three ways to support continuous data profiling: (1) it automatically displays data distributions and summary statistics to facilitate data comprehension; (2) it is live, so visualizations are always accessible and update automatically as the data updates; (3) it supports follow up analysis and documentation by authoring code for the user in the notebook. In a user study with 16 participants, we evaluate two versions of our system that integrate different levels of automation: both automatically show data profiles and facilitate code authoring, however, one version updates reactively ("live") and the other updates only on demand ("dead"). We find that both tools, dead or alive, facilitate insight discovery with 91% of user-generated insights originating from the tools rather than manual profiling code written by users. Participants found live updates intuitive and felt it helped them verify their transformations while those with on-demand profiles liked the ability to look at past visualizations. We also present a longitudinal case study on how AutoProfiler helped domain scientists find serendipitous insights about their data through automatic, live data profiles. Our results have implications for the design of future tools that offer automated data analysis support.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dfdeb43-88bd-40f4-95f8-0a57e87286ffCited by top-tier papers8
- Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMsHuichen Will Wang, Larry Birnbaum, Vidya SetlurCHI 2025 · 11 citations
- Loops: Leveraging Provenance and Visualization to Support Exploratory Data Analysis in NotebooksKlaus Eckelt, Kiran Gadhave, Alexander Lex, Marc StreitIEEE VIS 2024 · 7 citations
- Divisi: Interactive Search and Visualization for Scalable Exploratory Subgroup AnalysisVenkatesh Sivaraman, Zexuan Li, Adam PererCHI 2025 · 5 citations
- Crowdsourced Think-Aloud StudiesZach Cutler, Lane Harrison, Carolina Nobre, Alexander LexCHI 2025 · 4 citations
- Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling ScriptsYunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang et al.CHI 2025 · 4 citations
Builds on7
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 114 citations
- B2: Bridging Code and Interactive Visualization in Computational NotebooksYifan Wu, Joseph M. Hellerstein, Arvind SatyanarayanUIST 2020 · 71 citations
- mage: Fluid Moves Between Code and Graphical Work in Computational NotebooksMary Beth Kery, Donghao Ren, Fred Hohman, Dominik Moritz et al.UIST 2020 · 66 citations
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark et al.VLDB 2022 · 61 citations
Related papers
- PipelineProfiler: A Visual Analytics Tool for the Exploration of AutoML PipelinesJorge Piazentin Ono, Sonia Castelo, Roque Lopez, Enrico Bertini et al.IEEE VIS 2020 · 49 citations
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 25 citations
- Notable: On-the-fly Assistant for Data Storytelling in Computational NotebooksHaotian Li, Lu Ying, Haidong Zhang, Yingcai Wu et al.CHI 2023 · 38 citations
- NoteFlow: Leveraging Charts as Sight Glasses for Consistent and Continuous Data Flow TracingYuan Tian, Dazhen Deng, Sen Yang, Huawei Zheng et al.CHI 2026 · 1 citation
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 34 citations
