AerialVLA: A Vision-Language-Action Model for Aerial Navigation with Online Dialogue
Jinyu Chen, Hongyu Li, Zongheng Tang, Xiaoduo Li, Wenjun Wu, Si Liu
Abstract
Visual Dialogue Navigation (VDN) aims to enable agents to reach target locations through dialogue with humans. The integration of VDN into Unmanned Aerial Vehicle (UAV) systems enhances human-machine interaction by enabling intuitive, hands-free operation, thereby unlocking vast applications. However, existing VDN models for UAVs can only perform navigation based on dialogue history, lacking proactive interaction capabilities to correct trajectories. Moreover, their sequential observation history recording mechanism struggles to accurately localize landmarks observed in the historical context, leading to ineffective utilization of referential information in new user instructions. To address these, we present AerialVLA, an end-to-end UAV navigation framework integrating dialogue comprehension, action decision-making, and navigational question generation. Aeri-alVLA comprises three core components: i) we propose the Progress-Driven Navigation-Query Alternation mechanism to determine optimal questioning timing through navigation progress estimation autonomously. ii) To effectively model long-horizon history observation sequences, we develop the History Spatial-Temporal Fusion module that extracts discriminative spatial-temporal representations from historical observations. iii) Furthermore, to overcome data scarcity in training, we devise the Online Task-Driven Augmentation strategy that enhances learning through action-conditioned data augmentation. Experimental results demonstrate that AerialVLA achieves state-of-the-art navigation performance while exhibiting effective dialogue capabilities. Moreover, to better evaluate the agent's proactive dialogue and navigation abilities, our evaluation benchmark, named UAV Navigation with Online Dialogue (UNOD), incorporates an online dialogue interaction module. The UNOD assesses UAV agents' real-time questioning capabilities by leveraging an Air Commander Large Language Model to simulate human-UAV interactions during testing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23ecc51b-876d-4279-87a4-49e8539314d3Builds on10
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie et al.EMNLP 2020 · 208 citations
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang et al.ICCV 2023 · 132 citations
- HOP: History-and-Order Aware Pretraining for Vision-and-Language NavigationYanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu et al.CVPR 2022 · 71 citations
- Self-Motivated Communication Agent for Real-World Vision-Dialog NavigationYi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang et al.ICCV 2021 · 41 citations
Related papers
- Learning Fine-Grained Alignment for Aerial Vision-Dialog NavigationYifei Su, Dong An, Kehan Chen, Weichen Yu et al.AAAI 2025 · 7 citations
- AeroDuo: Aerial Duo for UAV-based Vision and Language NavigationRuipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang et al.ACM MM 2025 · 6 citations
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language NavigationXichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang et al.AAAI 2026 · 3 citations
- Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction ParsingShanshan Li, Jiawei Hou, Da Huang, Yanwei Fu et al.ACM MM 2025
