RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
Mingfei Han, Liang Ma, Kamila Zhumakhanova, Ekaterina Radionova, Jingyi Zhang, Xiaojun Chang, Xiaodan Liang, Ivan Laptev
Abstract
Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address this, we introduce RoomTour3D, a video-instruction dataset derived from web-based room tour videos that capture real-world indoor spaces and human walking demonstrations. Unlike existing VLN datasets, RoomTour3D leverages the scale and diversity of online videos to generate open-ended human walking trajectories and open-world navigable instructions. To compensate for the lack of navigation data in online videos, we perform 3D reconstruction and obtain 3D trajectories of walking paths augmented with additional information on the room types, object locations and 3D shape of surrounding scenes. Our dataset includes →100K open-ended description-enriched trajectories with →200K instructions, and 17K actionenriched trajectories from 1847 room tour environments. We demonstrate experimentally that RoomTour3D enables significant improvements across multiple VLN tasks including CVDN, SOON, R2R, and REVERIE. Moreover, Room-Tour3D facilitates the development of trainable zero-shot VLN agents, showcasing the potential and challenges of advancing towards open-world navigation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma et al.CVPR 2026 · 92 citations
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language NavigationShuo Wang, Yongcai Wang, Wanting Li, Xudong Cai et al.NeurIPS 2025 · 28 citations
- IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor EnvironmentsXu Liu, Yu Liu, Hanshuo Qiu, Qirong Yang et al.AAAI 2026 · 5 citations
- Lifelong Embodied Navigation LearningXudong Wang, Jiahua Dong, Baichen Liu, Qi Lyu et al.ICLR 2026 · 5 citations
- Lifting Unlabeled Internet-level Data for 3D Scene UnderstandingYixin Chen, Yaowei Zhang, Huangyue Yu, Junchao He et al.CVPR 2026 · 1 citation
Builds on30
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
Related papers
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li et al.ICCV 2023 · 57 citations
- VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationJialu Li, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit BansalAAAI 2024 · 13 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
- A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation LearningAishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh et al.CVPR 2023
- Curriculum Learning for Vision-and-Language NavigationJiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie PengNeurIPS 2021 · 33 citations
