Generating Landmark Navigation Instructions from Maps as a Graph-to-Text Problem
Raphael Schumann, Stefan Riezler
摘要
Car-focused navigation services are based on turns and distances of named streets, whereas navigation instructions naturally used by humans are centered around physical objects called landmarks. We present a neural model that takes OpenStreetMap representations as input and learns to generate navigation instructions that contain visible and salient landmarks from human natural language instructions. Routes on the map are encoded in a location-and rotation-invariant graph representation that is decoded into natural language instructions. Our work is based on a novel dataset of 7,672 crowd-sourced instances that have been verified by human navigation in Street View. Our evaluation shows that the navigation instructions generated by our system have similar properties as human-generated instructions, and lead to successful human navigation in Street View.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Evaluating the World Model Implicit in a Generative ModelKeyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon M. Kleinberg 等NeurIPS 2024 · 被引用 166 次
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
- Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor AreasRaphael Schumann, Stefan RiezlerACL 2022 · 被引用 38 次
- CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global MemoryWeichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng 等ACL 2025 · 被引用 22 次
- CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?Siqi Wang, Chao Liang, Yunfan Gao, Erxin Yu 等ICLR 2026 · 被引用 8 次
它引用的顶会 Paper2
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath 等AAAI 2020 · 被引用 78 次
相关 Paper
- VLN-Trans: Translator for the Vision and Language Navigation AgentYue Zhang, Parisa KordjamshidiACL 2023 · 被引用 6 次
- Scene Map-based Prompt Tuning for Navigation Instruction GenerationSheng Fan, Rui Liu, Wenguan Wang, Yi YangCVPR 2025
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationJungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto 等ICCV 2025 · 被引用 8 次
- Topological Planning With Transformers for Vision-and-Language NavigationKevin Chen, Junshen K. Chen, Jo Chuang, Marynel Vázquez 等CVPR 2021
- Less is More: Generating Grounded Navigation Instructions from LandmarksSu Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar 等CVPR 2022 · 被引用 41 次
