AURA: Multi-modal Shared Autonomy for Urban Navigation
Yukai Ma, Honglin He, Selina Song, Wayne Wu, Bolei Zhou
Abstract
Long-horizon navigation in complex urban environments still relies heavily on continuous human operation, which leads to fatigue, reduced efficiency, and safety concerns. Shared autonomy, where a Vision-Language AI agent and a human operator collaborate on maneuvering the mobile machine, presents a promising solution to address these issues. However, existing shared autonomy methods often require humans and AI to operate in the same action space, resulting in high cognitive overhead. We present Assistive Urban Robot Autonomy (AURA), a new multi-modal framework that decomposes urban navigation into high-level human instruction and low-level AI control. AURA incorporates a Spatial-Aware Instruction Encoder to align human instructions with visual and spatial context. To facilitate training, we construct UrbanWalks, a large-scale dataset composed of teleoperation and vision-language description data. Experiments in simulation and the real world demonstrate that AURA effectively follows human instructions, reduces manual operation effort, and improves navigation stability, while enabling online adaptation and continuous learning.Moreover, under similar takeover conditions, our hierarchical shared autonomy framework reduces human operation Frequency by over 75%. Code and data will be made available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 287c18c0-614f-4f09-9cf5-da5bedad7334Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human TrajectoriesYanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang et al.AAAI 2026
- AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsSudipta Paul, Amit Roy-Chowdhury, Anoop CherianNeurIPS 2022 · 43 citations
- Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction ParsingShanshan Li, Jiawei Hou, Da Huang, Yanwei Fu et al.ACM MM 2025
- UrbanLLaVA: A Multi-Modal Large Language Model for Urban IntelligenceJie Feng, Shengyuan Wang, Tianhui Liu, Yanxin Xi et al.ICCV 2025 · 7 citations
