SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
Gengze Zhou, Yicong Hong, Zun Wang, Chongyang Zhao, Mohit Bansal, Qi Wu
Abstract
The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in which the former emphasizes the exploration process, while the latter concentrates on following detailed textual commands. Despite the differing focuses of these tasks, the underlying requirements of interpreting instructions, comprehending the surroundings, and inferring action decisions remain consistent. This paper consolidates diverse navigation tasks into a unified and generic framework -- we investigate the core difficulties of sharing general knowledge and exploiting task-specific capabilities in learning navigation and propose a novel State-Adaptive Mixture of Experts (SAME) model that effectively enables an agent to infer decisions based on different-granularity language and dynamic observations. Powered by SAME, we present a versatile agent capable of addressing seven navigation tasks simultaneously that outperforms or achieves highly comparable performance to task-specific agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e004160d-9a34-4d30-993f-efcfbf8d16f9Cited by top-tier papers4
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
- Lifelong Embodied Navigation LearningXudong Wang, Jiahua Dong, Baichen Liu, Qi Lyu et al.ICLR 2026 · 5 citations
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao et al.CVPR 2026 · 4 citations
- Bootstrapping Language-Guided Navigation Learning with Self-Refining Data FlywheelZun Wang, Jialu Li, Yicong Hong, Songze Li et al.ICLR 2025
Builds on57
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
Related papers
- Towards Versatile Embodied NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangNeurIPS 2022 · 48 citations
- ULN: Towards Underspecified Vision-and-Language NavigationWeixi Feng, Tsu-Jui Fu, Yujie Lu, William Yang WangEMNLP 2022 · 2 citations
- ME: Continual Vision-and-Language Navigation via Mixture of Macro and Micro ExpertsYongliang Jiang, Huaidong Zhang, Xuandi Luo, Shengfeng HeICLR 2026
- LANA: A Language-Capable Navigator for Instruction Following and GenerationXiaohan Wang, Wenguan Wang, Jiayi Shao, Yi YangCVPR 2023
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu et al.ICCV 2025 · 1 citation
