Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I. Hong, Adam Perer
摘要
Machine learning models with high accuracy on test data can still produce systematic failures, such as harmful biases and safety issues, when deployed in the real world. To detect and mitigate such failures, practitioners run behavioral evaluation of their models, checking model outputs for specific types of inputs. Behavioral evaluation is important but challenging, requiring that practitioners discover real-world patterns and validate systematic failures. We conducted 18 semi-structured interviews with ML practitioners to better understand the challenges of behavioral evaluation and found that it is a collaborative, use-case-first process that is not adequately supported by existing task- and domain-specific tools. Using these findings, we designed zeno, a general-purpose framework for visualizing and testing AI systems across diverse use cases. In four case studies with participants using zeno on real-world models, we found that practitioners were able to reproduce previous manual analyses and discover new systematic failures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
- TeachTune: Reviewing Pedagogical Agents Against Diverse Student Profiles with Simulated StudentsHyoungwook Jin, Minju Yoo, Jeongeon Park, Yokyung Lee 等CHI 2025 · 被引用 58 次
- Interactive Debugging and Steering of Multi-Agent AI SystemsWill Epperson, Gagan Bansal, Victor C. Dibia, Adam Fourney 等CHI 2025 · 被引用 33 次
- Canvil: Designerly Adaptation for LLM-Powered User ExperiencesK. J. Kevin Feng, Q. Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan 等CHI 2025 · 被引用 14 次
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression ExperimentsAngie W. Boggust, Venkatesh Sivaraman, Yannick Assogba, Donghao Ren 等IEEE VIS 2024 · 被引用 12 次
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
- Domino: Discovering Systematic Errors with Cross-Modal EmbeddingsSabri Eyuboglu, Maya Varma, Khaled Kamal Saab, Jean-Benoit Delbrouck 等ICLR 2022 · 被引用 178 次
- DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative ModelsZijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang 等ACL 2023 · 被引用 149 次
相关 Paper
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 被引用 60 次
- Planning for Natural Language Failures with the AI PlaybookMatthew K. Hong, Adam Fourney, Derek DeBellis, Saleema AmershiCHI 2021 · 被引用 54 次
- Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and NeedsAgathe Balayn, Natasa Rikalo, Jie Yang, Alessandro BozzonCHI 2023 · 被引用 10 次
- Analyzing Collaborative Challenges and Needs of UX Practitioners when Designing with AI/MLMeena Devii Muralikumar, David W. McDonaldCSCW 2024 · 被引用 6 次
- Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual ProgrammingRuofei Du, Na Li, Jing Jin, Michelle Carney 等CHI 2023 · 被引用 33 次
