DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, Heung-Yeung Shum
Abstract
We present DINO (DETR with Improved deNoising anchOr boxes), a state-of-the-art end-to-end object detector. % in this paper. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a mixed query selection method for anchor initialization, and a look forward twice scheme for box prediction. DINO achieves AP in epochs and AP in epochs on COCO with a ResNet-50 backbone and multi-scale features, yielding a significant improvement of AP and AP, respectively, compared to DN-DETR, the previous best DETR-like model. DINO scales well in both model size and data size. Without bells and whistles, after pre-training on the Objects365 dataset with a SwinL backbone, DINO obtains the best results on both COCO val2017 (AP) and test-dev (AP). Compared to other models on the leaderboard, DINO significantly reduces its model size and pre-training data size while achieving better results. Our code will be available at https://github.com/IDEACVR/DINO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c20cfc7-d8c4-4110-8286-02f394a1fefcCited by top-tier papers230
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 313 citations
- Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision ApplicationsYuwen Xiong, Zhiqi Li, Yuntao Chen, Feng Wang et al.CVPR 2024 · 205 citations
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan et al.NeurIPS 2023 · 138 citations
- FocalFormer3D : Focusing on Hard Instance for 3D Object DetectionYilun Chen, Zhiding Yu, Yukang Chen, Shiyi Lan et al.ICCV 2023 · 109 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and SegmentationFeng Li, Hao Zhang, Huaizhe Xu, Shilong Liu et al.CVPR 2023
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 594 citations
- Detection Transformer with Stable MatchingShilong Liu, Tianhe Ren, Jiayu Chen, Zhaoyang Zeng et al.ICCV 2023 · 62 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
