MaDS: Long-Horizon GUI Automation via Synergizing Dual-Layer Memory and Multi-Round Debate
Pengchen Chen, Shi Chen, Qiming Ye, Xinli Chen, Xinran Li, Wei Xiang
摘要
Automating Graphical User Interface (GUI) operations with Multimodal Large Language Models (MLLMs) is promising but remains bottlenecked in real-world long-horizon settings. Key challenges include ensuring precise grounding across diverse interfaces and handling irreversible errors in extended workflows. Current methods often struggle to distinguish targets in low Signal-to-Noise Ratio (SNR) environments and lack sufficient preexecution verification to prevent error accumulation. To address this, we propose the Memory-augmented Debate System (MaDS). Specifically, MaDS combines: (1) a Dual-Layer Memory Module that integrates universal interaction priors with scenario-specific operational experience to mitigate grounding hallucinations; and (2) Multi-Round Debate that performs pre-execution verification, while transforming execution failures into retrievable Negative Warnings to reduce repeated errors. Additionally, we introduce MaDS-Benchmark, a benchmark for long-horizon mobile GUI tasks with process-oriented evaluation. Experiments show that MaDS achieves a 90.23% Task Success Rate on MaDS-Benchmark and strong performance on public benchmarks including AITW, AITZ, CAGUI, and GUIOdyssey.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White 等CHI 2021 · 被引用 145 次
- GTA1: GUI Test-time Scaling AgentYan Yang, Dongxu Li, Yutong Dai, Yuhao Yang 等ICLR 2026 · 被引用 109 次
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao 等MobiCom 2024 · 被引用 94 次
- Multi-Object Hallucination in Vision Language ModelsXuweiyi Chen, Ziqiao Ma, Xuejun Zhang, Sihan Xu 等NeurIPS 2024 · 被引用 77 次
相关 Paper
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language ModelsSitong Fang, Shiyi Hou, Kaile Wang, Boyuan Chen 等ICML 2026
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI AgentBin Kang, Shaoguo Wen, Yifei Bi, Shunlong Wu 等ICLR 2026
- RedDebate: Safer Responses Through Multi-Agent Red Teaming DebatesAli Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan ZhuICML 2026 · 被引用 7 次
- WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop EnvironmentsHaoren Zhao, Tianyi Chen, Zhen WangICML 2026
