NDH-Full: Learning and Evaluating Navigational Agents on Full-Length Dialogue
Hyounghun Kim, Jialu Li, Mohit Bansal
摘要
Communication between human and mobile agents is getting increasingly important as such agents are widely deployed in our daily lives. Vision-and-Dialogue Navigation is one of the tasks that evaluate the agent's ability to interact with humans for assistance and navigate based on natural language responses. In this paper, we explore the Navigation from Dialogue History (NDH) task, which is based on the Cooperative Vision-and-Dialogue Navigation (CVDN) dataset, and present a stateof-the-art model which is built upon Vision-Language transformers. However, despite achieving competitive performance, we find that the agent in the NDH task is not evaluated appropriately by the primary metric -Goal Progress. By analyzing the performance mismatch between Goal Progress and other metrics (e.g., normalized Dynamic Time Warping) from our state-of-the-art model, we show that NDH's sub-path based task setup (i.e., navigating partial trajectory based on its correspondent subset of the full dialogue) does not provide the agent with enough supervision signal towards the goal region. Therefore, we propose a new task setup called NDH-FULL which takes the full dialogue and the whole navigation path as one instance. We present a strong baseline model and show initial results on this new task. We further describe several approaches that we try, in order to improve the model performance (based on curriculum learning, pre-training, and data-augmentation), suggesting potential useful training methods on this new NDH-FULL task. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang 等ICCV 2023 · 被引用 136 次
- PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationJialu Li, Mohit BansalNeurIPS 2023 · 被引用 110 次
- Envedit: Environment Editing for Vision-and-Language NavigationJialu Li, Hao Tan, Mohit BansalCVPR 2022 · 被引用 76 次
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 被引用 13 次
- Bootstrapping Language-Guided Navigation Learning with Self-Refining Data FlywheelZun Wang, Jialu Li, Yicong Hong, Songze Li 等ICLR 2025
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby StepsWang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng 等ACL 2020 · 被引用 62 次
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
相关 Paper
- Vision-Dialog Navigation by Exploring Cross-Modal MemoryYi Zhu, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin 等CVPR 2020
- Improving Vision-and-Language Navigation by Generating Future-View Image SemanticsJialu Li, Mohit BansalCVPR 2023
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 被引用 33 次
- Learning Fine-Grained Alignment for Aerial Vision-Dialog NavigationYifei Su, Dong An, Kehan Chen, Weichen Yu 等AAAI 2025 · 被引用 7 次
- Tree-Structured Trajectory Encoding for Vision-and-Language NavigationXinzhe Zhou, Yadong MuAAAI 2023 · 被引用 2 次
