Understanding and Visualizing Data Iteration in Machine Learning
Fred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur Patel
摘要
Successful machine learning (ML) applications require iterations on both modeling and the underlying data. While prior visualization tools for ML primarily focus on modeling, our interviews with 23 ML practitioners reveal that they improve model performance frequently by iterating on their data (e.g., collecting new data, adding labels) rather than their models. We also identify common types of data iterations and associated analysis tasks and challenges. To help attribute data iterations to model performance, we design a collection of interactive visualizations and integrate them into a prototype, CHAMELEON, that lets users compare data features, training/testing splits, and performance across data versions. We present two case studies where developers apply CHAMELEON to their own evolving datasets on production ML projects. Our interface helps them verify data collection efforts, find failure cases stretching across data versions, capture data processing changes that impacted performance, and identify opportunities for future data iterations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Whither AutoML? Understanding the Role of Automation in Machine Learning WorkflowsDoris Xin, Eva Yiwei Wu, Doris Jung Lin Lee, Niloufar Salehi 等CHI 2021 · 被引用 103 次
- Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output LabelsJochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat 等CHI 2022 · 被引用 70 次
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach 等CSCW 2022 · 被引用 58 次
- The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive VisualizationMd. Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton 等CHI 2024 · 被引用 48 次
相关 Paper
- Symphony: Composing Interactive Interfaces for Machine LearningAlex Bäuerle, Ángel Alexander Cabrera, Fred Hohman, Megan Maher 等CHI 2022 · 被引用 47 次
- Diff in the Loop: Supporting Data Comparison in Exploratory Data AnalysisApril Yi Wang, Will Epperson, Robert A. DeLine, Steven Mark DruckerCHI 2022 · 被引用 34 次
- Angler: Helping Machine Translation Practitioners Prioritize Model ImprovementsSamantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery 等CHI 2023 · 被引用 20 次
- VIME: Visual Interactive Model Explorer for Identifying Capabilities and Limitations of Machine Learning Models for Sequential Decision-MakingAnindya Das Antar, Somayeh Molaei, Yan-Ying Chen, Matthew L. Lee 等UIST 2024 · 被引用 3 次
- Marcelle: Composing Interactive Machine Learning Workflows and InterfacesJules Françoise, Baptiste Caramiaux, Téo SanchezUIST 2021 · 被引用 37 次
