ViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced Network
Zhuochen Yu, Bijie Qiu, Andy W. H. Khong
Abstract
The sparsity of point clouds and inadequacy of semantic information pose challenges to current LiDAR-only 3D object detection methods. Recent methods alleviate these challenges by converting RGB images into virtual points via depth completion to be fused with LiDAR points. Although these methods have shown outstanding results, they often introduce significant computation overhead due to the high density of virtual points and noise due to inaccurate depth completion. Besides, they do not thoroughly leverage semantic information from images. In this work, we propose the virtual key instance enhanced network (ViKIENet), a highly efficient and effective multi-modal feature fusion framework that fuses the features of virtual key instances (VKIs) and LiDAR points through multiple stages. Our contributions include three main components: semantic key instance selection (SKIS), virtual-instance-focused fusion (VIFF), and virtual-instance-to-real attention (VIRA). We also propose the extended version ViKIENet-R with VIFF-R which includes rotationally equivariant features. Experiment results show that ViKIENet and ViKIENet-R achieve significant improvements in detection performance on the KITTI, JRDB, and nuScenes datasets compared to existing works. On the KITTI dataset, ViKIENet and ViKIENet-R operate at 22.7 and 15.0 FPS, respectively. As of CVPR submission (Nov. 15 th , 2024), ViKIENet ranks first on the car detection and orientation estimation leaderboard, while ViKIENet-R ranks second (compared with officially published papers) on the 3D car detection leaderboard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- Multimodal Virtual Point 3D DetectionTianwei Yin, Xingyi Zhou, Philipp KrähenbühlNeurIPS 2021 · 379 citations
- Focal Sparse Convolutional Networks for 3D Object DetectionYukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun et al.CVPR 2022 · 293 citations
- Sparse Fuse Dense: Towards High Quality 3D Detection with Depth CompletionXiaopei Wu, Liang Peng, Honghui Yang, Liang Xie et al.CVPR 2022 · 248 citations
Related papers
- Virtual Sparse Convolution for Multimodal 3D Object DetectionHai Wu, Chenglu Wen, Shaoshuai Shi, Xin Li et al.CVPR 2023
- Sparse Query Dense: Enhancing 3D Object Detection with Pseudo PointsYujian Mo, Yan Wu, Junqiao Zhao, Zhenjie Hou et al.ACM MM 2024 · 17 citations
- LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal FusionXin Li, Tao Ma, Yuenan Hou, Botian Shi et al.CVPR 2023
- Paint and Distill: Boosting 3D Object Detection with Semantic Passing NetworkBo Ju, Zhikang Zou, Xiaoqing Ye, Minyue Jiang et al.ACM MM 2022 · 12 citations
- PVGNet: A Bottom-Up One-Stage 3D Object Detector With Integrated Multi-Level FeaturesZhenwei Miao, Jikai Chen, Hongyu Pan, Ruiwen Zhang et al.CVPR 2021
