CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, Chen Feng
Abstract
https://ai4ce.github.io/CityWalker/ Crossing CityWalker Turn Sign Obstacle Traffic Light Road blocked Dense Traffic Proximity Web-scale Videos (2000+ hours) Expert Data (6 hours) Figure 1. Embodied Urban Navigation. Navigating urban spaces is challenging for (especially off-street) mobile agents. The differently colored pins ( ) along the route highlight various critical scenarios unique to complex and dynamic urban landscapes. Thumbnails on the right with corresponding colored pins demonstrate the real-world observation of these challenging cases. Our CityWalker model is trained with over 2000 hours of city walking videos and fine-tuned with a small amount of expert data to address these challenges effectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 043daff4-6ae3-4d4f-9253-28c32280d37cCited by top-tier papers13
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao et al.ICLR 2026 · 51 citations
- Astra: General Interactive World Model with Autoregressive DenoisingYixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao et al.ICLR 2026 · 29 citations
- SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied NavigationZiyi Chen, Yingnan Guo, Zedong Chu, Minghua Luo et al.CVPR 2026 · 19 citations
- CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local NavigationKai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu et al.ICLR 2026 · 12 citations
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor NavigationXia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma et al.CVPR 2026 · 8 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
Related papers
- UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human TrajectoriesYanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang et al.AAAI 2026
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour VideosMingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu et al.ICLR 2026 · 8 citations
- UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban SpacesBaining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang et al.ACL 2025 · 31 citations
- CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban EnvironmentsHaotian Xu, Yue Hu, Zhengqiu Zhu, Chen Gao et al.ACL 2026 · 6 citations
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li et al.ICLR 2026 · 93 citations
