When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks
Jieyu Zhou, Aryan Roy, Sneh Gupta, Daniel Weitekamp III, Christopher J. MacLellan
Abstract
Existing AI agents typically execute multi-step tasks autonomously and only allow user confirmation at the end. During execution, users have little control, making the confirm-at-end approach brittle: a single error can cascade and force a complete restart. Confirming every step avoids such failures, but imposes tedious overhead. Balancing excessive interruptions against costly rollbacks remains an open challenge. We address this problem by modeling confirmation as a minimum time scheduling problem. We conducted a formative study with eight participants, which revealed a recurring Confirmation-Diagnosis-Correction-Redo (CDCR) pattern in how users monitor errors. Based on this pattern, we developed a decision-theoretic model to determine time-efficient confirmation point placement. We then evaluated our approach using a within-subjects study where 48 participants monitored AI agents and repaired their mistakes while executing tasks. Results show that 81 percent of participants preferred our intermediate confirmation approach over the confirm-at-end approach used by existing systems, and task completion time was reduced by 13.54 percent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2ac2319-e24b-4d17-8ffe-a69c95fccb9dBuilds on40
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
Related papers
- Structure Enables Effective Self-Localization of Errors in LLMsAnkur Samanta, Akshayaa Magesh, Ayush Jain, Kavosh Asadi et al.ICML 2026 · 4 citations
- "What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step ProcessingJohannes Kirmayr, Raphael Paul Wennmacher, Khanh Huynh, Lukas Stappen et al.CHI 2026 · 1 citation
- Value of Information: A Framework for Human-Agent CommunicationYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang et al.ACL 2026 · 8 citations
- Interactive Debugging and Steering of Multi-Agent AI SystemsWill Epperson, Gagan Bansal, Victor C. Dibia, Adam Fourney et al.CHI 2025 · 33 citations
- Exploring the Design Space of User-System Communication for Smart-home Routine AssistantsYi-Shyuan Chiang, Ruei-Che Chang, Yi-Lin Chuang, Shih-Ya Chou et al.CHI 2020 · 35 citations
