CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark
Tianhang Wang, Xinhai Li, Fan Lu, Tianshi Gong, Jiankun Dong, Weiyi Xue, Sanqing Qu, Chenjia Bai, Guang Chen
Abstract
Vision-and-Language Navigation (VLN) primarily focuses on a single-agent-centric approach that executes human instructions step-by-step. In real environments with high demand or parallel workflows, collaboration VLN offers distinct benefits including shorter makespan and greater robustness through parallelism and role specialization. Collaboration VLN also brings new challenges including congestion, handoff errors, and rendezvous timing, which single-agent formulations overlook. Current datasets and protocols remain single-agent centered, which hides opportunities for assistance and ignores inter-robot interference. We fill this gap with Collaborative Long-Horizon VLN benchmark (CoNavBench), consisting of 4048 single and collaborative episodes with graph-level annotations and a collaboration type taxonomy that controls handoff styles and rendezvous patterns. To generate and evaluate at scale, we build NavCraft, an automated graph-grounded data generation platform. A two-stage hierarchical agent first produces a long-horizon base mission for the primary robot and then instantiates helper robots, allocates subgoals, and specifies validated handoffs and rendezvous. The agents operate with a scene graph in the loop derived from Habitat-Sim, which enables reachability checks, travel time, and interference assessment, and iterative schedule repair via an efficiency tool library. As a reference, we provide a collaborative baseline based on a finetuned Qwen2.5-VL-3B. Trained with CoNavBench, collaborative policies reduce makespan and improve reliability over strong single robot counterparts, yielding 18.11% step level success. Anonymous Website: https://navcraft.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c6ddf35-d36d-4a32-9f67-7b3843e4d293Builds on19
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility GenerationJiawei Wang, Renhe Jiang, Chuang Yang, Zengqing Wu et al.NeurIPS 2024 · 181 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
Related papers
- IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor EnvironmentsXu Liu, Yu Liu, Hanshuo Qiu, Qirong Yang et al.AAAI 2026 · 5 citations
- Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and MethodologyXiangyu Wang, Donglin Yang, Ziqin Wang, Hohin Kwan et al.ICLR 2025
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and MethodXinshuai Song, Weixing Chen, Yang Liu, Weikai Chen et al.CVPR 2025
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationYash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang et al.ICLR 2026 · 22 citations
