VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
Yiwei Zhang, Jin Gao, Fudong Ge, Guan Luo, Bing Li, Zhao-Xiang Zhang, Haibin Ling, Weiming Hu
摘要
Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, generating the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. The question is how to align the PV features with the generative models to facilitate the map estimation. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on https://github.com/Z1zyw/VQ-Map.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OmniGen: Unified Multimodal Sensor Generation for Autonomous DrivingTao Tang, Enhui Ma, Xia Zhou, Letian Wang 等ACM MM 2025 · 被引用 1 次
- HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving ScenesFudong Ge, Jin Gao, Hanshi Wang, Yiwei Zhang 等AAAI 2026
- ARINBEV: Bird's-Eye View Layout Estimation with Conditional Autoregressive ModelJiyong Kwag, Charles K. Toth, Alper YilmazICLR 2026
它引用的顶会 Paper20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- MapPrior: Bird's-Eye View Map Layout Estimation with Generative ModelsXiyue Zhu, Vlas Zyrianov, Zhijian Liu, Shenlong WangICCV 2023 · 被引用 20 次
- Improving Bird's Eye View Semantic Segmentation by Task DecompositionTianhao Zhao, Yongcan Chen, Yu Wu, Tianyang Liu 等CVPR 2024 · 被引用 11 次
- MapUQ: Map with Uncertainty Quantification for Robust BEV Vectorized ConstructionShaoyuan Mo, MaQi, l r, Bohan Li 等ICML 2026
- BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware RasterizationYixin Xiong, Ke Wang, Tongtong Cheng, Chunhui Liu 等CVPR 2026
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
