SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects
Abhinav Kumar, Yuliang Guo, Xinyu Huang, Liu Ren, Xiaoming Liu
Abstract
Improve KITTI-360 Val SoTA. (b) Improve nuScenes Val SoTA. (c) Theory Advancement. Figure 1. Teaser (a) SoTA frontal detectors struggle with large objects (low APLrg) even on a nearly balanced KITTI-360 dataset (Skewness in Fig. 7). Our proposed SeaBird achieves significant Mono3D improvements, particularly for large objects. (b) SeaBird also improves two SoTA BEV detectors, BEVerse-S [116] and HoP [121] on the nuScenes dataset, particularly for large objects. (c) Plot of convergence variance Var(ϵ) of dice and regression losses with the noise σ in depth prediction. The y-axis denotes the deviation from the optimal weight, so the lower the better. SeaBird leverages dice loss, which we prove is more noise-robust than regression losses for large objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd8009f2-5916-417e-8cf4-fb9058d5c5cbCited by top-tier papers6
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 13 citations
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour VideosMingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu et al.ICLR 2026 · 8 citations
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
- OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous DrivingMingqian Ji, Shanshan Zhang, Jian YangICCV 2025 · 2 citations
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang et al.ICCV 2025 · 1 citation
Builds on57
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
Related papers
- BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective SupervisionChenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao et al.CVPR 2023
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D Object Detection Network with Sufficient Depth CluesJinrang Jia, Zhenjia Li, Yifeng ShiNeurIPS 2023 · 69 citations
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie et al.ICCV 2023 · 65 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
