VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object Detection
Zihua Liu, Hiroki Sakuma, Masatoshi Okutomi
摘要
Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently illposed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels, typically obtained through expensive and laborintensive annotation on LiDAR point clouds. To tackle this problem, we propose a novel weakly supervised 3D object detection framework named VSRD (Volumetric Silhouette Rendering for Detection) to train 3D object detectors without any 3D supervision but only weak 2D supervision. VSRD consists of multi-view 3D auto-labeling and subsequent training of monocular 3D object detectors using the pseudo labels generated in the auto-labeling stage. In the auto-labeling stage, we represent the surface of each instance as a signed distance field (SDF) and render its silhouette as an instance mask through our proposed instanceaware volumetric silhouette rendering. To directly optimize the 3D bounding boxes through rendering, we decompose the SDF of each instance into the SDF of a cuboid and the residual distance field (RDF) that represents the residual from the cuboid. This mechanism enables us to optimize the 3D bounding boxes in an end-to-end manner by comparing the rendered instance masks with the ground truth instance masks. The optimized 3D bounding boxes serve as effective training data for 3D object detection. We conduct extensive experiments on the KITTI-360 dataset, demonstrating that our method outperforms the existing weakly supervised 3D object detection methods. The code is available at https://github.com/skmhrk1209/VSRD . * Equal contribution. The order was determined by a coin flip.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 被引用 3 次
- Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language GuidanceMingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang 等ICCV 2025
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman 等ICCV 2021 · 被引用 2,700 次
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt 等NeurIPS 2021 · 被引用 2,500 次
相关 Paper
- WeakM3D: Towards Weakly Supervised Monocular 3D Object DetectionLiang Peng, Senbo Yan, Boxi Wu, Zheng Yang 等ICLR 2022 · 被引用 25 次
- Weakly Supervised 3D Object Detection from Point CloudsZengyi Qin, Jinglu Wang, Yan LuACM MM 2020 · 被引用 68 次
- Weakly Supervised Monocular 3D Detection with a Single-View ImageXueying Jiang, Sheng Jin, Lewei Lu, Xiaoqin Zhang 等CVPR 2024
- IDA-3D: Instance-Depth-Aware 3D Object Detection From Stereo Vision for Autonomous DrivingWanli Peng, Hao Pan, He Liu, Yi SunCVPR 2020
- MonoNeRD: NeRF-like Representations for Monocular 3D Object DetectionJunkai Xu, Liang Peng, Haoran Chen, Hao Li 等ICCV 2023 · 被引用 54 次
