MonoCLUE: Object-Aware Clustering Enhances Monocular 3D Object Detection
Sunghun Yang, Minhyeok Lee, Jungho Lee, Sangyoun Lee
Abstract
Monocular 3D object detection offers a cost-effective solution for autonomous driving, but it suffers from the ill-posed depth and a limited field of view. These constraints lead to the lack of geometric cues and reduced accuracy in occluded or truncated scenes. While recent approaches incorporate additional depth information to address geometric ambiguity, they overlook the importance of visual cues essential for robust object recognition. In this paper, we propose MonoCLUE that enhances monocular 3D detection by leveraging both local clustering and generalized scene memory of visual features. First, we perform K-means clustering on visual features to capture distinct object-level appearance visual parts (e.g., bonnet, car roof), which improves the detection of partially visible objects. The clustered features are then propagated across the entire region to capture objects with similar appearances. Second, we construct a generalized scene memory by aggregating clustered features across images, providing consistent appearance representations that generalize scenes. This improves the consistency of object-level features, enabling stable detection across varying environments. Lastly, we integrate both local cluster features and generalized scene memory into object queries, guiding attention toward informative regions in the feature map. Exploiting an unified local clustering and generalized scene memory strategy, MonoCLUE enables robust monocular 3D detection under occlusion and limited visibility. Our proposed model achieves state-of-the-art performance on the KITTI benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 13 citations
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang et al.ICCV 2019 · 339 citations
Related papers
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- MonoCD: Monocular 3D Object Detection with Complementary DepthsLongfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang et al.CVPR 2024 · 52 citations
- MonoPair: Monocular 3D Object Detection Using Pairwise Spatial RelationshipsYongjian Chen, Lei Tai, Kai Sun, Mingyang LiCVPR 2020
- Homography Loss for Monocular 3D Object DetectionJiaqi Gu, Bojian Wu, Lubin Fan, Jianqiang Huang et al.CVPR 2022 · 47 citations
- Difficulty-Aware Label-Guided Denoising for Monocular 3D Object DetectionSoyul Lee, Seungmin Baek, Dongbo MinAAAI 2026
