MonoNeRD: NeRF-like Representations for Monocular 3D Object Detection
Junkai Xu, Liang Peng, Haoran Chen, Hao Li, Wei Qian, Ke Li, Wenxiao Wang, Deng Cai
摘要
In the field of monocular 3D detection, it is common practice to utilize scene geometric clues to enhance the detector's performance. However, many existing works adopt these clues explicitly such as estimating a depth map and back-projecting it into 3D space. This explicit methodology induces sparsity in 3D representations due to the increased dimensionality from 2D to 3D, and leads to substantial information loss, especially for distant and occluded objects. To alleviate this issue, we propose MonoNeRD, a novel detection framework that can infer dense 3D geometry and occupancy. Specifically, we model scenes with Signed Distance Functions (SDF), facilitating the production of dense 3D representations. We treat these representations as Neural Radiance Fields (NeRF) and then employ volume rendering to recover RGB images and depth maps. To the best of our knowledge, this work is the first to introduce volume rendering for M3D, and demonstrates the potential of implicit reconstruction for image-based 3D perception. Extensive experiments conducted on the KITTI-3D benchmark and Waymo Open Dataset demonstrate the effectiveness of MonoNeRD. Codes are available at https: //github.com/cskkxjk/MonoNeRD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 31 次
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu 等CVPR 2024 · 被引用 31 次
- MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion ModelsYasiru Ranasinghe, Deepti Hegde, Vishal M. PatelCVPR 2024 · 被引用 21 次
- Multi-View Attentive Contextualization for Multi-View 3D Object DetectionXianpeng Liu, Ce Zheng, Ming Qian, Nan Xue 等CVPR 2024 · 被引用 5 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
它引用的顶会 Paper33
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman 等ICCV 2021 · 被引用 2,700 次
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua 等NeurIPS 2020 · 被引用 1,535 次
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 被引用 1,421 次
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- PlenOctrees for Real-time Rendering of Neural Radiance FieldsAlex Yu, Ruilong Li, Matthew Tancik, Hao Li 等ICCV 2021 · 被引用 1,284 次
相关 Paper
- VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object DetectionZihua Liu, Hiroki Sakuma, Masatoshi OkutomiCVPR 2024
- NeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene ReconstructionWenyuan Zhang, Emily Yue-ting Jia, Junsheng Zhou, Baorui Ma 等CVPR 2025
- VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View ReconstructionYufan Ren, Fangjinhua Wang, Tong Zhang, Marc Pollefeys 等CVPR 2023
- H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in MotionHongyi Xu, Thiemo Alldieck, Cristian SminchisescuNeurIPS 2021 · 被引用 225 次
- Neural RGB-D Surface ReconstructionDejan Azinovic, Ricardo Martin-Brualla, Dan B. Goldman, Matthias Nießner 等CVPR 2022 · 被引用 272 次
