GeniNav: Generative Model Driven Image-Goal Navigation via Imagination-Guided Consistency Flow Matching
Yuqi Chen, Junjie Gao, Yongzhou Pan, Siyuan Song, ZIXUAN ZHANG, Jiaping Xiao, Mir Feroskhan
Abstract
Image-goal navigation driven by generative models has recently shown strong potential owing to their ability to perform multi-modal reasoning and stable learning in continuous control spaces. Despite their promise, current methods still face several fundamental limitations. Many rely on pre-built priors and lack explicit mechanisms for trajectory evaluation, restricting generalization and goal alignment in map-free navigation. Moreover, current generative policies often face inefficiency or temporal inconsistency, resulting in temporally unstable motion. The absence of interactive, closed-loop benchmarks further limits fair and reproducible comparison. To address these issues, we propose GeniNav, a generative image-goal navigation framework that couples a VLM-driven latent subgoal imagination module for high-level semantic guidance with Multi-Segment Consistency Flow Matching (MS-CFM) for temporally smooth and dynamically coherent motion generation. A hybrid trajectory evaluation module further integrates semantic alignment and geometric feasibility to assess goal consistency. We also introduce a closed-loop simulation benchmark with a large-scale dataset spanning 176 scenes and 491.6 km for standardized training and evaluation. Extensive experiments in simulation and on real robots demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67bab914-441b-4efb-ae03-eb4fd6ba2ba7Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Pathdreamer: A World Model for Indoor NavigationJing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge et al.ICCV 2021 · 128 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
Related papers
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationBadi Li, Renjie Lu, Yu Zhou, Jingke Meng et al.NeurIPS 2025 · 5 citations
- ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene ImaginationXinxin Zhao, Wenzhe Cai, Likun Tang, Teng WangICLR 2025
- MVGBench: A Comprehensive Benchmark for Multi-View Generation ModelsXianghui Xie, Jan Eric Lenssen, Gerard Pons-MollICCV 2025 · 2 citations
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality GenerationDogyun Park, Taehoon Lee, Minseok Joo, Hyunwoo J. KimNeurIPS 2025 · 4 citations
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingYang Zhou, Hao Shao, Letian Wang, Zhuofan Zong et al.ICLR 2026 · 20 citations
