CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang, Liu Ren, Xiaoming Liu
摘要
Monocular 3D object detectors, while effective on data from one ego camera height, struggle with unseen or out-of-distribution camera heights. Existing methods often rely on Plucker embeddings, image transformations or data augmentation. This paper takes a step towards this understudied problem by first investigating the impact of camera height variations on state-of-the-art (SoTA) Mono3D models. With a systematic analysis on the extended CARLA dataset with multiple camera heights, we observe that depth estimation is a primary factor influencing performance under height variations. We mathematically prove and also empirically observe consistent negative and positive trends in mean depth error of regressed and ground-based depth models, respectively, under camera height changes. To mitigate this, we propose Camera Height Robust Monocular 3D Detector (CHARM3R), which averages both depth estimates within the model. CHARM3R improves generalization to unseen camera heights by more than , achieving SoTA performance on the CARLA dataset. Codes and Models at https://github.com/abhi1kumar/CHARM3R
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 被引用 13 次
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain 等CVPR 2026 · 被引用 7 次
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi 等CVPR 2026 · 被引用 5 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
- CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object DetectionZhaonian Kuang, Rui Ding, Haotian Wang, Xinhu Zheng 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper67
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 被引用 542 次
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 被引用 527 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li 等ICCV 2021 · 被引用 404 次
相关 Paper
- Monocular 3D Object Detection: An Extrinsic Parameter Free ApproachYunsong Zhou, Yuan He, Hongzi Zhu, Cheng Wang 等CVPR 2021
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- BEVHeight: A Robust Framework for Vision-based Roadside 3D Object DetectionLei Yang, Kaicheng Yu, Tao Tang, Jun Li 等CVPR 2023
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionYung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li 等ICCV 2025 · 被引用 2 次
