ImVoteNet: Boosting 3D Object Detection in Point Clouds With Image Votes
Charles R. Qi, Xinlei Chen, Or Litany, Leonidas J. Guibas
摘要
3D object detection has seen quick progress thanks to advances in deep learning on point clouds. A few recent works have even shown state-of-the-art performance with just point clouds input (e.g. VOTENET). However, point cloud data have inherent limitations. They are sparse, lack color information and often suffer from sensor noise. Images, on the other hand, have high resolution and rich texture. Thus they can complement the 3D geometry provided by point clouds. Yet how to effectively use image information to assist point cloud based detection is still an open question. In this work, we build on top of VOTENET and propose a 3D detection architecture called IMVOTENET specialized for RGB-D scenes. IMVOTENET is based on fusing 2D votes in images and 3D votes in point clouds. Compared to prior work on multi-modal detection, we explicitly extract both geometric and semantic features from the 2D images. We leverage camera parameters to lift these features to 3D. To improve the synergy of 2D-3D feature fusion, we also propose a multi-tower training scheme. We validate our model on the challenging SUN RGB-D dataset, advancing state-of-the-art results by 5.7 mAP. We also provide rich ablation studies to analyze the contribution of each design choice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu 等ICCV 2021 · 被引用 368 次
- DeepInteraction: 3D Object Detection via Modality InteractionZeyu Yang, Jiaqi Chen, Zhenwei Miao, Wei Li 等NeurIPS 2022 · 被引用 268 次
- Surface Representation for Point CloudsHaoxi Ran, Jun Liu, Chengjie WangCVPR 2022 · 被引用 230 次
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 214 次
它引用的顶会 Paper1
相关 Paper
- ImOV3D: Learning Open Vocabulary Point Clouds 3D Object Detection from Only 2D ImagesTiming Yang, Yuanliang Ju, Li YiNeurIPS 2024 · 被引用 22 次
- RBGNet: Ray-based Grouping for 3D Object DetectionHaiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang 等CVPR 2022 · 被引用 63 次
- MLCVNet: Multi-Level Context VoteNet for 3D Object DetectionQian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang 等CVPR 2020
- ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object DetectionTao Tu, Shun-Po Chuang, Yu-Lun Liu, Cheng Sun 等ICCV 2023 · 被引用 16 次
- Bridged Transformer for Vision and Point Cloud 3D Object DetectionYikai Wang, TengQi Ye, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 55 次
