Understanding and Visualizing Data Iteration in Machine Learning
Fred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur Patel
Abstract
Successful machine learning (ML) applications require iterations on both modeling and the underlying data. While prior visualization tools for ML primarily focus on modeling, our interviews with 23 ML practitioners reveal that they improve model performance frequently by iterating on their data (e.g., collecting new data, adding labels) rather than their models. We also identify common types of data iterations and associated analysis tasks and challenges. To help attribute data iterations to model performance, we design a collection of interactive visualizations and integrate them into a prototype, CHAMELEON, that lets users compare data features, training/testing splits, and performance across data versions. We present two case studies where developers apply CHAMELEON to their own evolving datasets on production ML projects. Our interface helps them verify data collection efforts, find failure cases stretching across data versions, capture data processing changes that impacted performance, and identify opportunities for future data iterations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Whither AutoML? Understanding the Role of Automation in Machine Learning WorkflowsDoris Xin, Eva Yiwei Wu, Doris Jung Lin Lee, Niloufar Salehi et al.CHI 2021 · 103 citations
- Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output LabelsJochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat et al.CHI 2022 · 70 citations
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach et al.CSCW 2022 · 58 citations
- The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive VisualizationMd. Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton et al.CHI 2024 · 48 citations
Related papers
- Symphony: Composing Interactive Interfaces for Machine LearningAlex Bäuerle, Ángel Alexander Cabrera, Fred Hohman, Megan Maher et al.CHI 2022 · 47 citations
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 34 citations
- Angler: Helping Machine Translation Practitioners Prioritize Model ImprovementsSamantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery et al.CHI 2023 · 20 citations
- VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-MakingAnindya Das Antar, Somayeh Molaei, Yan-Ying Chen, Matthew L. Lee et al.UIST 2024 · 3 citations
- Marcelle: Composing Interactive Machine Learning Workflows and InterfacesJules Françoise, Baptiste Caramiaux, Téo SanchezUIST 2021 · 37 citations
