Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas
Raphael Schumann, Stefan Riezler
摘要
Vision and language navigation (VLN) is a challenging visually-grounded language understanding task. Given a natural language navigation instruction, a visual agent interacts with a graph-based environment equipped with panorama images and tries to follow the described route. Most prior work has been conducted in indoor scenarios where best results were obtained for navigation on routes that are similar to the training routes, with sharp drops in performance when testing on unseen environments. We focus on VLN in outdoor scenarios and find that in contrast to indoor VLN, most of the gain in outdoor VLN on unseen data is due to features like junction type embedding or heading delta that are specific to the respective environment graph, while image information plays a very minor role in generalizing VLN to unseen outdoor areas. These findings show a bias to specifics of graph representations of urban environments, demanding that VLN tasks grow in scale and diversity of geographical environments. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Evaluating the World Model Implicit in a Generative ModelKeyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg 等NeurIPS 2024 · 被引用 166 次
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li 等ICCV 2023 · 被引用 57 次
- Policy Improvement using Language Feedback ModelsVictor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre CôtéNeurIPS 2024 · 被引用 18 次
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 被引用 13 次
它引用的顶会 Paper4
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath 等AAAI 2020 · 被引用 78 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- Generating Landmark Navigation Instructions from Maps as a Graph-to-Text ProblemRaphael Schumann, Stefan RiezlerACL 2021
相关 Paper
- Loc4Plan: Locating Before Planning for Outdoor Vision and Language NavigationHuilin Tian, Jingke Meng, Wei-Shi Zheng, Yuan-Ming Li 等ACM MM 2024 · 被引用 6 次
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 被引用 29 次
- ULN: Towards Underspecified Vision-and-Language NavigationWeixi Feng, Tsu-Jui Fu, Yujie Lu, William Yang WangEMNLP 2022 · 被引用 2 次
- Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsHaodong Hong, Sen Wang, Zi Huang, Qi Wu 等ACM MM 2024 · 被引用 4 次
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
