Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical Voxelization
Qi Chen, Lin Sun, Ernest Cheung, Alan L. Yuille
Abstract
Recent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a.k.a. perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage the benefits from both BEV and RV. The widely-used cuboid-shaped voxels in Cartesian coordinate system only benefit BEV feature map. Therefore, to enable learning both BEV and RV feature maps, we introduce Hybrid-Cylindrical-Spherical voxelization. Our findings show that simply adding detection on another view as auxiliary supervision will lead to poor performance. We proposed a pair of cross-view transformers to transform the feature maps into the other view and introduce cross-view consistency loss on them. Comprehensive experiments on the challenging NuScenes Dataset validate the effectiveness of our proposed method which leverages joint optimization and complementary information on both views. Remarkably, our approach achieved mAP of 55.8%, outperforming all published approaches by at least 3% in overall performance and up to 16.5% in safety-crucial categories like cyclist. * indicates equal contributions 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a87f0a2c-483a-47a1-93a0-c87f1ef396f0Cited by top-tier papers29
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object DetectionYingwei Li, Adams Wei Yu, Tianjian Meng, Benjamin Caine et al.CVPR 2022 · 508 citations
- Multimodal Virtual Point 3D DetectionTianwei Yin, Xingyi Zhou, Philipp KrähenbühlNeurIPS 2021 · 379 citations
- Focal Sparse Convolutional Networks for 3D Object DetectionYukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun et al.CVPR 2022 · 293 citations
Builds on6
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
- Fast Point R-CNNYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaICCV 2019 · 440 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic SegmentationYang Zhang, Zixiang Zhou, Philip David, Xiangyu Yue et al.CVPR 2020
- What You See is What You Get: Exploiting Visibility for 3D Object DetectionPeiyun Hu, Jason Ziglar, David Held, Deva RamananCVPR 2020
Related papers
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
- Fully Convolutional One-Stage 3D Object Detection on LiDAR Range ImagesZhi Tian, Xiangxiang Chu, Xiaoming Wang, Xiaolin Wei et al.NeurIPS 2022 · 168 citations
- CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird’s-Eye-View Semantic SegmentationJeongbin Hong, Dooseop Choi, Taeg-Hyun An, KYOUNG AN AN et al.CVPR 2026 · 1 citation
- VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial AttentionShengheng Deng, Zhihao Liang, Lin Sun, Kui JiaCVPR 2022 · 92 citations
