NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
Zhihao Luo, Wentao Yan, Jingyu Gong, Min Wang, Zhizhong Zhang, Xuhong Wang, Yuan Xie, Xin Tan
Abstract
Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms. In this paper, we observe that both tasks can be formulated as Markov Decision Processes (MDP), suggesting a foundational principle for their unification. Hence, we present NaviMaster, the first unified agent capable of unifying GUI navigation and embodied navigation within a single framework. Specifically, NaviMaster (i) proposes a visual-target trajectory collection pipeline that generates trajectories for both GUI and embodied tasks using a single formulation. (ii) employs a unified reinforcement learning framework on the mix data to improve generalization. (iii) designs a novel distance-aware reward to ensure efficient learning from the trajectories. Through extensive experiments on out-of-domain benchmarks, NaviMaster is shown to outperform state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation. Ablation studies further demonstrate the efficacy of our unified training strategy, data mixing strategy, and reward design. Our codes, data, and checkpoints are available at https:// iron-boyy.github.io/navimaster-page/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f214e6e-efdc-4e9e-b530-046bad4c633dCited by top-tier papers3
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li et al.CVPR 2026 · 11 citations
- Experience-driven Multi-turn Reinforcement Learning for GUI AgentsZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen et al.ACL 2026
- GUI-SAGE: Enhancing GUI Automation with Self-Explanatory LearningFei Tang, Zhangxuan Gu, Zhengxi Lu, Shangzhan Zhang et al.CVPR 2026
Builds on14
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view InputQijian Tian, Xin Tan, Yuan Xie, Lizhuang MaAAAI 2025 · 45 citations
- FastLGS: Speeding Up Language Embedded Gaussians with Feature Grid MappingYuzhou Ji, He Zhu, Junshu Tang, Wuyi Liu et al.AAAI 2025 · 29 citations
- Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied ExplorationSen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma et al.CVPR 2026 · 14 citations
Related papers
- NavA³: Understanding Any Instruction, Navigating Anywhere, Finding AnythingLingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu et al.ACL 2026 · 24 citations
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
- From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsAndrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev et al.CVPR 2025
- MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationPingrui Zhang, Xianqiang Gao, Yuhan Wu, Kehui Liu et al.ICCV 2025
- ENTL: Embodied Navigation Trajectory LearnerKlemen Kotar, Aaron Walsman, Roozbeh MottaghiICCV 2023 · 16 citations
