Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory
Yu Qi, Hongyu Li, Shaofei Huang, Tianrui Hui, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Si Liu, Meng Wang
摘要
In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address these issues, we propose PSC-AVDN, a training-free framework that tightly couples a three-stage Parsing-Search-Confirmation reasoning pipeline with a Structured Spatial Memory (SSM).The parsing stage uses an LLM to convert ambiguous dialogue instructions into stable geometric directional and destination cues.A Search Chain-of-Thought (S-CoT) then performs stepwise target exploration under high-altitude observations, and a Confirmation Chain-of-Thought (C-CoT) conducts fine-grained verification around candidate regions to resolve visual ambiguity.Meanwhile, SSM integrates three complementary sources of spatial cues, including multi-scale visual observation, spatial visual memory, and structured geometric memory to provide global spatial context and long-horizon consistency.Extensive experiments on ANDH and ANDH-Full show that PSC-AVDN establishes new state-of-the-art performance in the training-free setting, matching or surpassing several finetuned methods.Code will be publicly available at: https://github.com/QY6616/PSC-AVDN
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 被引用 228 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- SOAT: A Scene- and Object-Aware Transformer for Vision-and-Language NavigationAbhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee 等NeurIPS 2021 · 被引用 88 次
相关 Paper
- LookasideVLN: Direction-Aware Aerial Vision-and-Language NavigationYuwei Ning, Ganlong Zhao, Yipeng Qin, Si Liu 等CVPR 2026 · 被引用 6 次
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationXichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang 等AAAI 2026 · 被引用 3 次
- Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation AgentsTianyi Ma, Yue Zhang, Zehao Wang, Parisa KordjamshidiACL 2026 · 被引用 3 次
- AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online DialogueJinyu Chen, Hongyu Li, Zongheng Tang, Xiaoduo Li 等AAAI 2026
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang 等ICCV 2023 · 被引用 132 次
