Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented Dialogs
Toby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell, Brad A. Myers
Abstract
A major problem in task-oriented conversational agents is the lack of support for the repair of conversational breakdowns. Prior studies have shown that current repair strategies for these kinds of errors are often ineffective due to: (1) the lack of transparency about the state of the system's understanding of the user's utterance; and (2) the system's limited capabilities to understand the user's verbal attempts to repair natural language understanding errors. This paper introduces SOVITE, a new multi-modal (speech plus direct manipulation) interface that helps users discover, identify the causes of, and recover from conversational breakdowns using the resources of existing mobile app GUIs for grounding. SOVITE displays the system's understanding of user intents using GUI screenshots, allows users to refer to third-party apps and their GUI screens in conversations as inputs for intent disambiguation, and enables users to repair breakdowns using direct manipulation on these screenshots. The results from a remote user study with 10 users using SOVITE in 7 scenarios suggested that SOVITE's approach is usable and effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 205ddd8f-2a81-4aef-a8f8-3325e23196e7Cited by top-tier papers22
- StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental InvolvementZheng Zhang, Ying Xu, Yanhao Wang, Bingsheng Yao et al.CHI 2022 · 168 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- VISAR: A Human-AI Argumentative Writing Assistant with Visual Programming and Rapid Draft PrototypingZheng Zhang, Jie Gao, Ranjodh Singh Dhaliwal, Toby Jia-Jun LiUIST 2023 · 101 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
- On the Design of AI-powered Code Assistants for NotebooksAndrew M. McNutt, Chenglong Wang, Robert A. DeLine, Steven Mark DruckerCHI 2023 · 78 citations
Builds on3
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu et al.ICSE 2020 · 101 citations
- The Role of Conversational Grounding in Supporting Symbiosis Between People and Digital AssistantsJanghee Cho, Emilee J. RaderCSCW 2020 · 50 citations
- Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web ApplicationsRitam Jyoti Sarmah, Yunpeng Ding, Di Wang, Cheuk Yin Phipson Lee et al.UIST 2020 · 19 citations
Related papers
- Understanding is a Two-Way Street: User-Initiated Repair on Agent Responses and Hearing in Conversational InterfacesRobert J. Moore, Sungeun An, Olivia H. MarreseCSCW 2024 · 3 citations
- Agent-SAMA: State-Aware Mobile AssistantLinqiang Guo, Wei Liu, Yi Wen Heng, Tse-Hsun (Peter) Chen et al.AAAI 2026 · 2 citations
- Dango: A Mixed-Initiative Data Wrangling System using Large Language ModelWei-Hao Chen, Weixi Tong, Amanda Case, Tianyi ZhangCHI 2025 · 19 citations
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection BehaviorPenghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu et al.NeurIPS 2025 · 20 citations
- Diagnosing and Prioritizing Issues in Automated Order-Taking Systems: A Machine-Assisted Error Discovery ApproachMaeda F. Hanafi, Frederick Reiss, Yannis Katsis, Robert Moore et al.CHI 2025 · 3 citations
