UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
Han Xiao, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Hao He, Lue Fan, Liuyang Bian, Rui Hu, Liang Liu, Shuai Ren, Yafei Wen
Abstract
In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectively. The reward model, UI-Genie-RM, features an image-text interleaved architecture that efficiently pro- cesses historical context and unifies action-level and task-level rewards. To sup- port the training of UI-Genie-RM, we develop deliberately-designed data genera- tion strategies including rule-based verification, controlled trajectory corruption, and hard negative mining. To address the second challenge, a self-improvement pipeline progressively expands solvable complex GUI tasks by enhancing both the agent and reward models through reward-guided exploration and outcome verification in dynamic environments. For training the model, we generate UI- Genie-RM-517k and UI-Genie-Agent-16k, establishing the first reward-specific dataset for GUI agents while demonstrating high-quality synthetic trajectory gen- eration without manual annotation. Experimental results show that UI-Genie achieves state-of-the-art performance across multiple GUI agent benchmarks with three generations of data-model self-improvement. We open-source our complete framework implementation and generated datasets to facilitate further research in https://github.com/Euphoria16/UI-Genie.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5835e56f-59af-41b4-869e-131366c6f7b0Cited by top-tier papers4
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought ReasoningXinyan Chen, Renrui Zhang, Dongzhi Jiang, Aojun Zhou et al.NeurIPS 2025 · 54 citations
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic ModelsZhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang et al.CVPR 2026 · 6 citations
- Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple ActionsGuo Gan, Yuxuan Ding, Cong Chen, Yuwei Ren et al.ACL 2026 · 6 citations
Builds on13
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
- AlphaMath Almost Zero: Process Supervision without ProcessGuoxin Chen, Minpeng Liao, Chengxi Li, Kai FanNeurIPS 2024 · 219 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng et al.ACL 2025 · 71 citations
Related papers
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsBofei Zhang, Zirui Shang, Zhi Gao, Wang Zhang et al.AAAI 2026 · 26 citations
- M-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data MiningRui Lyu, Juncheng Mo, Tianyi Chu, Chen Rao et al.ICLR 2026
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent PretrainingWeimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue et al.ICML 2026 · 2 citations
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level FilteringYifei He, Pranit Chawla, Yaser Souri, Subhojit Som et al.ACL 2026 · 5 citations
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection BehaviorPenghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu et al.NeurIPS 2025 · 20 citations
