UniMODE: Unified Monocular 3D Object Detection
Zhuoling Li, Xiaogang Xu, Ser-Nam Lim, Hengshuang Zhao
摘要
Realizing unified monocular 3D object detection, including both indoor and outdoor scenes, holds great importance in applications like robot navigation. However, involving various scenarios of data to train models poses challenges due to their significantly different characteristics, e.g., diverse geometry properties and heterogeneous domain distributions. To address these challenges, we build a detector based on the bird's-eye-view (BEV) detection paradigm, where the explicit feature projection is beneficial to addressing the geometry learning ambiguity when employing multiple scenarios of data to train detectors. Then, we split the classical BEV detection architecture into two stages and propose an uneven BEV grid design to handle the convergence instability caused by the aforementioned challenges. Moreover, we develop a sparse BEV feature projection strategy to reduce computational cost and a unified domain alignment method to handle heterogeneous domains. Combining these techniques, a unified detector UniMODE is derived, which surpasses the previous state-of-the-art on the challenging Omni3D dataset (a large-scale dataset including both indoor and outdoor scenes) by 4.9% AP 3D , revealing the first successful generalization of a BEV detector to unified 3D object detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Generalizing Visual Geometry Priors to Sparse Gaussian Occupancy PredictionChangqing Zhou, Yueru Luo, Changhao ChenCVPR 2026 · 被引用 10 次
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich 等AAAI 2025 · 被引用 7 次
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun 等ICCV 2025 · 被引用 6 次
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-SightYunze Man, Shihao Wang, Guowen Zhang, Johan Bjorck 等CVPR 2026 · 被引用 6 次
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionYung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper17
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 被引用 542 次
相关 Paper
- Towards Generalizable Multi-Camera 3D Object Detection via Perspective RenderingHao Lu, Yunpeng Zhang, Guoqing Wang, Qing Lian 等AAAI 2025 · 被引用 5 次
- GPA-3D: Geometry-aware Prototype Alignment for Unsupervised Domain Adaptive 3D Object Detection from Point CloudsZiyu Li, Jingming Guo, Tongtong Cao, Bingbing Liu 等ICCV 2023 · 被引用 19 次
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao 等NeurIPS 2023 · 被引用 65 次
- One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object DetectionZhenyu Wang, Yali Li, Hengshuang Zhao, Shengjin WangNeurIPS 2024 · 被引用 13 次
- CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionGyusam Chang, Wonseok Roh, Sujin Jang, Dongwook Lee 等AAAI 2024 · 被引用 8 次
