Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
Yuxiang Zhou, Jichang Li, Yanhao Zhang, Haonan Lu, Guanbin Li
摘要
Mobile agents show immense potential, yet current state-of-the-art (SoTA) agents exhibit inadequate success rates on real-world, long-horizon, cross-application tasks. We attribute this bottleneck to the agents' excessive reliance on static, internal knowledge within MLLMs, which leads to two critical failure points: 1) strategic hallucinations in high-level planning and 2) operational errors during low-level execution on user interfaces (UI). The core insight of this paper is that high-level planning and low-level UI operations require fundamentally distinct types of knowledge. Planning demands high-level, strategy-oriented experiences, whereas operations necessitate low-level, precise instructions closely tied to specific app UIs. Motivated by these insights, we propose Mobile-Agent-RAG, a novel hierarchical multi-agent framework that innovatively integrates dual-level retrieval augmentation. At the planning stage, we introduce Manager-RAG to reduce strategic hallucinations by retrieving human-validated comprehensive task plans that provide high-level guidance. At the execution stage, we develop Operator-RAG to improve execution accuracy by retrieving the most precise low-level guidance for accurate atomic actions, aligned with the current app and subtask. To accurately deliver these knowledge types, we construct two specialized retrieval-oriented knowledge bases. Furthermore, we introduce Mobile-Eval-RAG, a challenging benchmark for evaluating such agents on realistic multi-app, long-horizon tasks. Extensive experiments demonstrate that Mobile-Agent-RAG significantly outperforms SoTA baselines, improving task completion rate by 11.0% and step efficiency by 10.2%, establishing a robust paradigm for context-aware, reliable multi-agent mobile automation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented GenerationZiyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang 等EMNLP 2025
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question AnsweringRong Cheng, Jinyi Liu, Yan Zheng, Fei Ni 等ACL 2025
- Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory GraphsZheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu 等EMNLP 2024 · 被引用 6 次
- EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic RetrievalJiashi Lin, Changhong Jiang, Xiangru Lin, Ruifei Zhang 等CVPR 2026 · 被引用 2 次
- InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task PlanningZheng Wang, Shu Xian Teo, Jun Jie Chew, Wei ShiSIGIR 2025 · 被引用 4 次
