Pix2Map: Cross-Modal Retrieval for Inferring Street Maps from Images
Xindi Wu, KwunFung Lau, Francesco Ferroni, Aljosa Osep, Deva Ramanan
摘要
Self-driving vehicles rely on urban street maps for autonomous navigation. In this paper, we introduce Pix2Map, a method for inferring urban street map topology directly from ego-view images, as needed to continually update and expand existing maps. This is a challenging task, as we need to infer a complex urban road topology directly from raw image data. The main insight of this paper is that this problem can be posed as cross-modal retrieval by learning a joint, cross-modal embedding space for images and existing maps, represented as discrete graphs that encode the topological layout of the visual surroundings. We conduct our experimental evaluation using the Argoverse dataset and show that it is indeed possible to accurately retrieve street maps corresponding to both seen and unseen roads solely from image data. Moreover, we show that our retrieved maps can be used to update or expand existing maps and even show proof-of-concept results for visual localization and image retrieval from spatial graphs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map ConstructionYi Zhou, Hui Zhang, Jiaqian Yu, Yifan Yang 等CVPR 2024 · 被引用 19 次
- Improving Online Lane Graph Extraction by Object-Lane ClusteringYigit Baran Can, Alexander Liniger, Danda Pani Paudel, Luc Van GoolICCV 2023 · 被引用 11 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas 等ICCV 2021 · 被引用 329 次
- CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional ConvolutionLizhe Liu, Xiaohao Chen, Siyu Zhu, Ping TanICCV 2021 · 被引用 312 次
- Topological Map Extraction From Overhead ImagesZuoyue Li, Jan Dirk Wegner, Aurélien LucchiICCV 2019 · 被引用 181 次
- Structured Bird's-Eye-View Traffic Scene Understanding from Onboard ImagesYigit Baran Can, Alexander Liniger, Danda Pani Paudel, Luc Van GoolICCV 2021 · 被引用 147 次
相关 Paper
- Learning Global Representation from Queries for Vectorized HD Map ConstructionShoumeng Qiu, Xinrun Li, Yang Long, Xiangyang Xue 等ICML 2026 · 被引用 1 次
- Beyond Single view Decoding: Dual-view Map Inference from Trajectories via Primal-Dual Graphs Co-generationWenyu Wu, Jiafan Liu, Jiali MaoWWW 2026
- VectorMapNet: End-to-end Vectorized HD Map LearningYicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang 等ICML 2023 · 被引用 332 次
- HDMapGen: A Hierarchical Graph Generative Model of High Definition MapsLu Mi, Hang Zhao, Charlie Nash, Xiaohan Jin 等CVPR 2021
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
