Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
Guo Gan, Yuxuan Ding, Cong Chen, Yuwei Ren, Yin Huang, Hong Zhou
Abstract
Online reinforcement learning (RL) serves as an effective method for enhancing the capabilities of Android agents. However, guiding agents to learn through online interaction is prohibitively expensive due to the high latency of emulators and the sample inefficiency of existing RL algorithms. We identify a fundamental limitation in current approaches: the Single State Single Action paradigm, which updates the policy with one-to-one state-action pairs from online one-way rollouts without fully exploring each costly emulator state. In this paper, we propose Android Coach, a novel framework that shifts the training paradigm to Single State Multiple Actions, allowing the agent to sample and utilize multiple actions for a single online state. We enable this without additional emulator overhead by learning a critic that estimates action values. To ensure the critic serves as a reliable coach, we integrate a process reward model and introduce a group-wise advantage estimator based on the averaged critic outputs. Extensive experiments demonstrate the effectiveness and efficiency of Android Coach: it achieves 7.5% and 8.3% success rate improvements on AndroidLab and AndroidWorld over UI-TARS-1.5-7B, and attains 1.4x higher training efficiency than Single State Single Action methods PPO and GRPO at matched success rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5d452ee-f25c-4a4e-a785-76d841a7e894Cited by top-tier papers3
- NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation TasksZhihao Luo, Wentao Yan, Jingyu Gong, Min Wang et al.ACL 2026 · 13 citations
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
- Anchor: Branch-Point Data Generation for GUI AgentsJinbiao Wei, Yilun Zhao, Kangqi Ni, Arman CohanACL 2026 · 3 citations
Builds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri et al.NeurIPS 2024 · 239 citations
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine et al.ICML 2024 · 163 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
Related papers
- Experience-driven Multi-turn Reinforcement Learning for GUI AgentsZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen et al.ACL 2026
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI AgentsYifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu et al.ICLR 2026 · 45 citations
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng et al.ACL 2025 · 71 citations
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin et al.AAAI 2026 · 103 citations
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMsYujie Zhao, Lanxiang Hu, Yang Wang, Minmin Hou et al.ICLR 2026 · 26 citations
