DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data Preparation
Arpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi, Jane Hoffswell, Shunan Guo, Eunyee Koh, Shamkant B. Navathe, Alex Endert
Abstract
Selecting relevant data subsets from large, unfamiliar datasets can be difficult. We address this challenge by modeling and visualizing two kinds of auxiliary information: (1) quality – the validity and appropriateness of data required to perform certain analytical tasks; and (2) usage – the historical utilization characteristics of data across multiple users. Through a design study with 14 data workers, we integrate this information into a visual data preparation and analysis tool, DataPilot. DataPilot presents visual cues about “the good, the bad, and the ugly” aspects of data and provides graphical user interface controls as interaction affordances, guiding users to perform subset selection. Through a study with 36 participants, we investigate how DataPilot helps users navigate a large, unfamiliar tabular dataset, prepare a relevant subset, and build a visualization dashboard. We find that users selected smaller, effective subsets with higher quality and usage, and with greater success and confidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a327d7ef-099c-4c9b-b589-beeba5423f7cCited by top-tier papers2
- StructVizor: Interactive Profiling of Semi-Structured Textual DataYanwei Huang, Yan Miao, Di Weng, Adam Perer et al.CHI 2025 · 3 citations
- Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeekYanwei Huang, Arpit NarechaniaCHI 2026 · 2 citations
Builds on9
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Expanding Explainability: Towards Social Transparency in AI systemsUpol Ehsan, Q. Vera Liao, Michael J. Muller, Mark O. Riedl et al.CHI 2021 · 505 citations
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 260 citations
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine et al.CHI 2020 · 110 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
Related papers
- Improving Visualization Interpretation Using CounterfactualsSmiti Kaul, David Borland, Nan Cao, David GotzIEEE VIS 2021 · 26 citations
- Selecting Sub-tables for Data ExplorationYael Amsterdamer, Susan B. Davidson, Tova Milo, Kathy Razmadze et al.ICDE 2023 · 5 citations
- "I Need to Find That One Chart": How Data Workers Navigate, Summarize and Communicate Analytical ConversationsKen Gu, Srishti Palani, Vidya SetlurCHI 2026 · 1 citation
- A Heuristic Approach for Dual Expert/End-User Evaluation of Guidance in Visual AnalyticsDavide Ceneda, Christopher Collins, Mennatallah El-Assady, Silvia Miksch et al.IEEE VIS 2023 · 9 citations
- How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz StudyKen Gu, Madeleine Grunde-McLaughlin, Andrew M. McNutt, Jeffrey Heer et al.CHI 2024 · 37 citations
