Angler: Helping Machine Translation Practitioners Prioritize Model Improvements
Samantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery, Fred Hohman
Abstract
Machine learning (ML) models can fail in unexpected ways in the real world, but not all model failures are equal. With finite time and resources, ML practitioners are forced to prioritize their model debugging and improvement efforts. Through interviews with 13 ML practitioners at Apple, we found that practitioners construct small targeted test sets to estimate an error’s nature, scope, and impact on users. We built on this insight in a case study with machine translation models, and developed Angler, an interactive visual analytics tool to help practitioners prioritize model improvements. In a user study with 7 machine translation experts, we used Angler to understand prioritization practices when the input space is infinite, and obtaining reliable signals of model quality is expensive. Our study revealed that participants could form more interesting and user-focused hypotheses for prioritization by analyzing quantitative summary statistics and qualitatively assessing data by reading sentences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e46a5ff-9628-4655-8b06-6570e916b7d5Cited by top-tier papers12
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
- Farsight: Fostering Responsible AI Awareness During AI Application PrototypingZijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry et al.CHI 2024 · 55 citations
- Canvil: Designerly Adaptation for LLM-Powered User ExperiencesK. J. Kevin Feng, Q. Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan et al.CHI 2025 · 14 citations
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression ExperimentsAngie W. Boggust, Venkatesh Sivaraman, Yannick Assogba, Donghao Ren et al.IEEE VIS 2024 · 12 citations
- Towards a Non-Ideal Methodological Framework for Responsible MLRamaravind Kommiya Mothilal, Shion Guha, Syed Ishtiaque AhmedCHI 2024 · 7 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Data Augmentation Can Improve RobustnessSylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg et al.NeurIPS 2021 · 427 citations
- Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic BehaviorsHong Shen, Alicia DeVos, Motahhare Eslami, Kenneth HolsteinCSCW 2021 · 156 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
Related papers
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 60 citations
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 114 citations
- Explaining mispredictions of machine learning models using rule inductionJürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali et al.FSE 2021 · 26 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
- Planning for Natural Language Failures with the AI PlaybookMatthew K. Hong, Adam Fourney, Derek DeBellis, Saleema AmershiCHI 2021 · 54 citations
