From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
Honglin He, Yukai Ma, Brad Squicciarini, Wayne Wu, Bolei Zhou
Abstract
Navigation foundation models trained on massive web-scale data enable agents to generalize across diverse environments and embodiments. However, these models, which are trained solely on offline data, often lack the capacity to reason about the consequences of their actions or adapt through counterfactual understanding. They thus face significant limitations in the real-world urban navigation where interactive and safe behaviors, such as avoiding obstacles and moving pedestrians, are critical. To tackle these challenges, we introduce the Seeing-to-Experiencing (S2E) learning framework to scale the capability of navigation foundation models with reinforcement learning. S2E combines the strengths of pre-training on offline videos and post-training through reinforcement learning. It maintains the model's generalizability acquired from large-scale real-world videos while enhancing its interactivity through reinforcement learning in simulation environments. Specifically, we introduce two innovations: 1) an Anchor-Guided Distribution Matching strategy for offline pretraining, which stabilizes learning and models diverse motion patterns through anchor-based supervision; and 2) a Residual-Attention Module for reinforcement learning, which obtains reactive behaviors from simulation environments without erasing the model's pretrained knowledge. Moreover, we establish a comprehensive end-to-end evaluation benchmark, NavBench-GS, built on photorealistic 3D Gaussian Splatting reconstructions of real-world scenes that incorporate physical interactions. It can systematically assess the generalizability and safety of navigation foundation models. Extensive experiments show that S2E mitigates the diminishing returns often seen when scaling with offline data alone. We perform a thorough analysis of the benefits of Reinforcement Learning (RL) compared to Supervised Fine-Tuning (SFT) in the context of post-training for robot learning. Our findings emphasize the crucial role of integrating interactive online experiences to effectively scale foundation models in Robotics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7191bd5-fa14-4c8e-86f0-4f31630d94a4Cited by top-tier papers4
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language NavigationMeng Wei, Chenyang Wan, Jiaqi Peng, Xiqian Yu et al.ICLR 2026 · 77 citations
- SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied NavigationZiyi Chen, Yingnan Guo, Zedong Chu, Minghua Luo et al.CVPR 2026 · 19 citations
- TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic EnvironmentsZhiyu Huang, Yun Zhang, Johnson Liu, Rui Song et al.ICML 2026 · 11 citations
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour VideosMingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu et al.ICLR 2026 · 8 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Self-Improving Embodied Foundation ModelsSeyed Kamyar Seyed Ghasemipour, Ayzaan Wahid, Jonathan Tompson, Pannag Sanketi et al.NeurIPS 2025 · 38 citations
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline MatchingHao Zhong, Muzhi Zhu, Shenyan Zeng, Anzhou Li et al.CVPR 2026 · 1 citation
- SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and CollaborationYan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen et al.NeurIPS 2025 · 9 citations
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang et al.CVPR 2026 · 40 citations
- CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local NavigationKai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu et al.ICLR 2026 · 12 citations
