Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction
Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou, Jiwen Lu
Abstract
Modern methods for vision-centric autonomous driving perception widely adopt the bird's-eye-view (BEV) representation to describe a 3D scene. Despite its better efficiency than voxel representation, it has difficulty describing the fine-grained 3D structure of a scene with a single plane. To address this, we propose a tri-perspective view (TPV) representation which accompanies BEV with two additional perpendicular planes. We model each point in the 3D space by summing its projected features on the three planes. To lift image features to the 3D TPV space, we further propose a transformer-based TPV encoder (TPVFormer) to obtain the TPV features effectively. We employ the attention mechanism to aggregate the image features corresponding to each query in each TPV plane. Experiments show that our model trained with sparse supervision effectively predicts the semantic occupancy for all voxels. We demonstrate for the first time that using only camera inputs can achieve comparable performance with LiDAR-based methods on the LiDAR segmentation task on nuScenes. Code: https://github.com/wzzheng/TPVFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b874f11b-cc02-4ac9-bf2d-3cb5a9a4e705Cited by top-tier papers157
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng et al.NeurIPS 2023 · 279 citations
- OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionXiaofeng Wang, Zheng Zhu, Wenbo Xu, Yunpeng Zhang et al.ICCV 2023 · 270 citations
- Scene as OccupancyWenwen Tong, Chonghao Sima, Tai Wang, Li Chen et al.ICCV 2023 · 251 citations
Builds on21
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPsChristian Reiser, Songyou Peng, Yiyi Liao, Andreas GeigerICCV 2021 · 963 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
Related papers
- COTR: Compact Occupancy TRansformer for Vision-Based 3D Occupancy PredictionQihang Ma, Xin Tan, Yanyun Qu, Lizhuang Ma et al.CVPR 2024 · 30 citations
- OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionZiteng Wen, Hai Xu, Chenyu Liu, Tao Guo et al.ACM MM 2023 · 5 citations
- TBP-Former: Learning Temporal Bird's-Eye-View Pyramid for Joint Perception and Prediction in Vision-Centric Autonomous DrivingShaoheng Fang, Zi Wang, Yiqi Zhong, Junhao Ge et al.CVPR 2023
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
- SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionPin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng et al.CVPR 2024 · 37 citations
