CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Sujay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, Chen Feng
摘要
https://ai4ce.github.io/CityWalker/ Crossing CityWalker Turn Sign Obstacle Traffic Light Road blocked Dense Traffic Proximity Web-scale Videos (2000+ hours) Expert Data (6 hours) Figure 1. Embodied Urban Navigation. Navigating urban spaces is challenging for (especially off-street) mobile agents. The differently colored pins ( ) along the route highlight various critical scenarios unique to complex and dynamic urban landscapes. Thumbnails on the right with corresponding colored pins demonstrate the real-world observation of these challenging cases. Our CityWalker model is trained with over 2000 hours of city walking videos and fine-tuned with a small amount of expert data to address these challenges effectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao 等ICLR 2026 · 被引用 51 次
- Astra: General Interactive World Model with Autoregressive DenoisingYixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao 等ICLR 2026 · 被引用 29 次
- SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied NavigationZiyi Chen, Yingnan Guo, Zedong Chu, Minghua Luo 等CVPR 2026 · 被引用 19 次
- CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local NavigationKai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu 等ICLR 2026 · 被引用 12 次
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor NavigationXia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma 等CVPR 2026 · 被引用 8 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
相关 Paper
- UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human TrajectoriesYanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang 等AAAI 2026
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour VideosMingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu 等ICLR 2026 · 被引用 8 次
- UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban SpacesBaining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang 等ACL 2025 · 被引用 31 次
- CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban EnvironmentsHaotian Xu, Yue Hu, Zhengqiu Zhu, Chen Gao 等ACL 2026 · 被引用 6 次
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li 等ICLR 2026 · 被引用 93 次
