Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal Navigation
Jiacong Zhou, Jiaxu Miao, Yourun Lin, Xianyun Wang, Jun Xiao, Jun Yu
摘要
Aerial object-goal navigation (Aerial ObjectNav) requires an Unmanned Aerial Vehicle (UAV) to navigate to target objects in large-scale outdoor environments using only visual observations and high-level object descriptions, without detailed step-by-step instructions. Existing approaches rely on local observations or short-term history, lacking comprehensive scene understanding and efficient spatial exploration strategies, which constrains their navigation capability in complex aerial scenarios. To address these challenges, we propose OctMem-Agent, an octree memory-augmented framework for aerial object-goal navigation. Specifically, we introduce an Adaptive Octree Memory that incrementally aggregates RGB-D observations into a hierarchical 3D representation, capturing both explored regions and unexplored frontiers across large-scale aerial environments. We further propose a Instruction-Guided Memory Query module that extracts task-relevant scene and exploration tokens through instruction-modulated queries. By integrating these tokens with visual observations and language instructions, OctoMem-Agent achieves comprehensive scene understanding and effective spatial exploration for target localization. Extensive experiments on the Aerial ObjectNav benchmark UAV-ON demonstrate that our method achieves a significant 7.5% improvement in success rate over existing methods, validating the effectiveness of our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
相关 Paper
- LookasideVLN: Direction-Aware Aerial Vision-and-Language NavigationYuwei Ning, Ganlong Zhao, Yipeng Qin, Si Liu 等CVPR 2026 · 被引用 6 次
- APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal NavigationDaoxuan Zhang, Ping Chen, Xiaobo Xia, Xiu Su 等CVPR 2026 · 被引用 10 次
- CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global MemoryWeichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng 等ACL 2025 · 被引用 22 次
- Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic MethodologyYatai Ji, Zhengqiu Zhu, Yong Zhao, Beidan Liu 等AAAI 2026 · 被引用 8 次
- Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and MethodologyXiangyu Wang, Donglin Yang, Ziqin Wang, Hohin Kwan 等ICLR 2025
