MLCVNet: Multi-Level Context VoteNet for 3D Object Detection
Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Yiming Zhang, Kai Xu, Jun Wang
Abstract
In this paper, we address the 3D object detection task by capturing multi-level contextual information with the selfattention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual information between these objects. Comparatively, we propose Multi-Level Context VoteNet (MLCVNet) to recognize 3D objects correlatively, building on the state-of-the-art VoteNet. We introduce three context modules into the voting and classifying stages of VoteNet to encode contextual information at different levels. Specifically, a Patch-to-Patch Context (PPC) module is employed to capture contextual information between the point patches, before voting for their corresponding object centroid points. Subsequently, an Object-to-Object Context (OOC) module is incorporated before the proposal and classification stage, to capture the contextual information between object candidates. Finally, a Global Scene Context (GSC) module is designed to learn the global scene context. We demonstrate these by capturing contextual information at patch, object and scene levels. Our method is an effective way to promote detection accuracy, achieving new state-of-the-art detection performance on challenging 3D object detection datasets, i.e., SUN RGBD and ScanNet. We also release our code at https://github.com/NUAAXQ/MLCVNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc68cb6e-cb88-494a-bab9-7c274af08368Cited by top-tier papers48
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 217 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point CloudsHaiyang Wang, Lihe Ding, Shaocong Dong, Shaoshuai Shi et al.NeurIPS 2022 · 110 citations
Builds on1
Related papers
- VENet: Voting Enhancement Network for 3D Object DetectionQian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang et al.ICCV 2021 · 60 citations
- Group Contextual Encoding for 3D Point CloudsXu Liu, Chengtao Li, Jian Wang, Jingbo Wang et al.NeurIPS 2020 · 5 citations
- A Hierarchical Graph Network for 3D Object Detection on Point CloudsJintai Chen, Biwen Lei, Qingyu Song, Haochao Ying et al.CVPR 2020
- RBGNet: Ray-based Grouping for 3D Object DetectionHaiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang et al.CVPR 2022 · 63 citations
- Back-Tracing Representative Points for Voting-Based 3D Object Detection in Point CloudsBowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang et al.CVPR 2021
