Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
Tianyi Ma, Yue Zhang, Zehao Wang, Parisa Kordjamshidi
摘要
Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3D environments. While recent progress has been driven by large-scale pre-training and data augmentation, current methods still struggle to generalize to unseen scenarios, particularly when complex spatial and temporal reasoning is required. In this work, we propose SkillNav, a modular framework that introduces structured, skill-based reasoning into Transformer-based VLN agents. Our method decomposes navigation into a set of interpretable atomic skills (e.g., Vertical Movement, Area and Region Identification, Stop and Pause), each handled by a specialized agent. To support targeted skill training without manual data annotation, we construct a synthetic dataset pipeline that generates diverse, linguistically natural, skill-specific instruction-trajectory pairs. We then introduce a novel training-free Vision-Language Model (VLM)-based router, which dynamically selects the most suitable agent at each time step by aligning sub-goals with visual observations and historical actions. SkillNav obtains competitive results on commonly used benchmarks and establishes state-of-the-art generalization to the GSA-R2R, a benchmark with novel instruction styles and unseen environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang 等ICCV 2023 · 被引用 136 次
- Vision-Language Navigation with Random Environmental MixupChong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang 等ICCV 2021 · 被引用 113 次
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 被引用 110 次
相关 Paper
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li 等CVPR 2026 · 被引用 7 次
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu 等CVPR 2026 · 被引用 10 次
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 被引用 33 次
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee 等NeurIPS 2021 · 被引用 88 次
- VLN-Trans: Translator for the Vision and Language Navigation AgentYue Zhang, Parisa KordjamshidiACL 2023 · 被引用 6 次
