CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
Zhuoyuan Yu, Yuxing Long, Zihan Yang, Chengyan Zeng, Hongwei Fan, Jiyao Zhang, Hao Dong
Abstract
PKU-Agibot Lab *Equal contribution, † Project Leader, ‡ Corresponding author https://correctnav.github.io Crowded Objects Avoidance Open-vocabulary Landmark Move Forward and Turn right at the human-like robot. Continue moving to stop near the yellow box. … Z-Shape Building Structure Walk down the corridor hallway in front of you and you will see an opened meeting room. Enter the … … … Walk straight along the hallway until you reach the red fire extinguisher box at the end and stop when you reach … Pedestrian Avoidance … Error Correction … in front of a white wall, turn right. Walk forward. When you see a green plant on your right front, stop. Drift Correction Walk straight and turn left in front of a wall. Walk straight and turn right at the opened door. Enter and walk to the wooden table. … … Walk until you reach the plant and turn left. Walk straight, turn left at the next corner, walk forward to the … Instruction Across Rooms Walk out of the kitchen room you are in and turn left. Move across the living room, walk to the end of the hallway and turn right .Walk into the bedroom and stop by the bed. Landmark State Change Move forward and turn right to walk through an opened doorway. … … … … … … C rrectNa Figure 1: Diverse Capabilities of CorrectNav. The model takes only monocular RGB video and language instructions as inputs, predicting navigation actions. Empowered by the Self-correction Flywheel post-training, CorrectNav not only maintains outstanding multimodal reasoning (Blue), but also displays improved deviation correction (Red), obstacle avoidance (Green), and complex action execution (Yellow).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao et al.ICLR 2026 · 51 citations
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation PredictionFei Liu, Shichao Xie, Minghua Luo, Zedong Chu et al.CVPR 2026 · 16 citations
- AwareVLN: Reasoning with Self-awareness for Vision-Language NavigationWenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li et al.CVPR 2026 · 7 citations
- DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language NavigationZihao Xin, Wentong Li, Yixuan Jiang, Bin Wang et al.CVPR 2026 · 6 citations
- AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language NavigationXin Ding, Jianyu Wei, Yifan Yang, Shiqi Jiang et al.ICML 2026 · 6 citations
Builds on16
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- Waypoint Models for Instruction-guided Navigation in Continuous EnvironmentsJacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee et al.ICCV 2021 · 153 citations
- Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng et al.NeurIPS 2022 · 143 citations
- Scaling Data Generation in Vision-and-Language NavigationZun Wang, Jialu Li, Yicong Hong, Yi Wang et al.ICCV 2023 · 136 citations
Related papers
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor NavigationXia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma et al.CVPR 2026 · 8 citations
- LANA: A Language-Capable Navigator for Instruction Following and GenerationXiaohan Wang, Wenguan Wang, Jiayi Shao, Yi YangCVPR 2023
- OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial ConstraintsMingjie Pan, Jiyao Zhang, Tianshu Wu, Yinghao Zhao et al.CVPR 2025
- Narrowing the Gap between Vision and Action in NavigationYue Zhang, Parisa KordjamshidiACM MM 2024 · 2 citations
- VLN-ChEnv: Vision-language Navigation in Changeable EnvironmentsShubo Liu, Hongsheng Zhang, Qian Qiao, Qi Wu et al.ACM MM 2025 · 2 citations
