Planning for Natural Language Failures with the AI Playbook
Matthew K. Hong, Adam Fourney, Derek DeBellis, Saleema Amershi
Abstract
Prototyping AI user experiences is challenging due in part to probabilistic AI models making it difficult to anticipate, test, and mitigate AI failures before deployment. In this work, we set out to support practitioners with early AI prototyping, with a focus on natural language (NL)-based technologies. Our interviews with 12 NL practitioners from a large technology company revealed that, in addition to challenges prototyping AI, prototyping was often not happening at all or focused only on idealized scenarios due to a lack of tools and tight timelines. These findings informed our design of the AI Playbook, an interactive and low-cost tool we developed to encourage proactive and systematic consideration of AI errors before deployment. Our evaluation of the AI Playbook demonstrates its potential to 1) encourage product teams to prioritize both ideal and failure scenarios, 2) standardize the articulation of AI failures from a user experience perspective, and 3) act as a boundary object between user experience designers, data scientists, and engineers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de380dc2-46df-4e3b-9946-671cf861063bCited by top-tier papers16
- Investigating How Practitioners Use Human-AI Guidelines: A Case Study on the People + AI GuidebookNur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg et al.CHI 2023 · 103 citations
- Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI ChallengesQiaosi Wang, Michael Madaio, Shaun K. Kane, Shivani Kapania et al.CHI 2023 · 90 citations
- Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User ExperienceQ. Vera Liao, Hariharan Subramonyam, Jennifer Wang, Jennifer Wortman VaughanCHI 2023 · 81 citations
- Farsight: Fostering Responsible AI Awareness During AI Application PrototypingZijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry et al.CHI 2024 · 55 citations
- Seamful XAI: Operationalizing Seamful Design in Explainable AIUpol Ehsan, Q. Vera Liao, Samir Passi, Mark O. Riedl et al.CSCW 2024 · 41 citations
Builds on2
- Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to DesignQian Yang, Aaron Steinfeld, Carolyn P. Rosé, John ZimmermanCHI 2020 · 604 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
Related papers
- fAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision TasksSteven Moore, Q. Vera Liao, Hariharan SubramonyamCHI 2023 · 36 citations
- Prototyping with Prompts: Emerging Approaches and Challenges in Generative AI Design for Collaborative Software TeamsHari Subramonyam, Divy Thakkar, Andrew Ku, Jürgen Dieber et al.CHI 2025 · 26 citations
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine LearningÁngel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein et al.CHI 2023 · 51 citations
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsGlen Berman, Nitesh Goyal, Michael MadaioCHI 2024 · 40 citations
- Tinker, Tailor, Configure, Customize: The Articulation Work of Contextualizing an AI Fairness ChecklistMichael A. Madaio, Jingya Chen, Hanna M. Wallach, Jennifer Wortman VaughanCSCW 2024 · 13 citations
