BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
Qinzhuo Wu, Pengzhi Gao, Wei Liu, Jian Luan
Abstract
Graphical User Interface (GUI) agents have gained substantial attention due to their impressive capabilities to complete tasks through multiple interactions within GUI environments. However, existing agents primarily focus on enhancing the accuracy of individual actions and often lack effective mechanisms for detecting and recovering from errors. To address these shortcomings, we propose the BacktrackAgent, a robust framework that incorporates a backtracking mechanism to improve task completion efficiency. BacktrackAgent includes verifier, judger, and reflector components as modules for error detection and recovery, while also applying judgment rewards to further enhance the agent's performance. Additionally, we develop a training dataset specifically designed for the backtracking mechanism, which considers the outcome pages after action executions. Experimental results show that BacktrackAgent has achieved performance improvements in both task success rate and step accuracy on Mobile3M and Auto-UI benchmarks. Our data and code will be released upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using AgentsBowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu et al.ACL 2026 · 15 citations
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI AutomationZehao Deng, Tianjie Ju, Zheng Wu, Zhuosheng Zhang et al.CVPR 2026 · 1 citation
- MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round DebatePengchen Chen, Shi Chen, Qiming Ye, Xinli Chen et al.ACL 2026
Builds on8
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov et al.ICML 2023 · 318 citations
- WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic ExplorationYao Zhang, Zijian Ma, Yunpu Ma, Zhen Han et al.AAAI 2025 · 101 citations
- Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM AgentsZhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo et al.ICML 2024 · 17 citations
Related papers
- M-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data MiningRui Lyu, Juncheng Mo, Tianyi Chu, Chen Rao et al.ICLR 2026
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection BehaviorPenghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu et al.NeurIPS 2025 · 20 citations
- COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI AutomationDi Zhao, Longhui Ma, Siwei Wang, Miao Wang et al.EMNLP 2025
- ProBench: Benchmarking GUI Agents with Accurate Process InformationLeyang Yang, Ziwei Wang, Xiaoxuan Tang, Sheng Zhou et al.AAAI 2026
- MMBench-GUI: A Unified Hierarchical Evaluation Framework for Multi-Platform GUI AgentsXuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding et al.CVPR 2026
