AlignDet: Aligning Pre-training and Fine-tuning in Object Detection
Ming Li, Jie Wu, Xionghui Wang, Chen Chen, Jie Qin, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan
Abstract
The paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limit the detector’s performance, generalization ability, and convergence speed. To this end, we propose AlignDet, a unified pre-training framework that can be adapted to various existing detectors to alleviate the discrepancies. AlignDet decouples the pre-training process into two stages, i.e., image-domain and box-domain pre-training. The image-domain pre-training optimizes the detection backbone to capture holistic visual abstraction, and box-domain pre-training learns instance-level semantics and task-aware concepts to initialize the parts out of the backbone. By incorporating the self-supervised pretrained backbones, we can pre-train all modules for various detectors in an unsupervised paradigm. As depicted in Figure 1, extensive experiments demonstrate that AlignDet can achieve significant improvements across diverse protocols, such as , and . For example, AlignDet improves FCOS by 5.3 mAP, RetinaNet by 2.1 mAP, Faster R-CNN by 3.3 mAP, and DETR by 2.3 mAP under fewer epochs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b5d6559-2863-47ed-9015-e1cd5c445dcaCited by top-tier papers3
- Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual ReconstructionChen-Long Duan, Yong Li, Xiu-Shen Wei, Lin ZhaoNeurIPS 2024 · 9 citations
- Bi-Level Optimization for Self-Supervised AI-Generated Face DetectionMian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan et al.ICCV 2025 · 2 citations
- DisCo DETR: Distance-aware Multi-view Contrastive Learning for DETR Pre-trainingChao Ouyang, Yuyang Bai, Jun Zhang, Tianlu Gao et al.AAAI 2026
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
Related papers
- DETReg: Unsupervised Pretraining with Region Priors for Object DetectionAmir Bar, Xin Wang, Vadim Kantorov, Colorado J. Reed et al.CVPR 2022 · 130 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningYizhou Wang, Meilin Chen, Shixiang Tang, Feng Zhu et al.NeurIPS 2022 · 2 citations
- PreDet: Large-scale weakly supervised pre-training for detectionVignesh Ramanathan, Rui Wang, Dhruv MahajanICCV 2021 · 14 citations
- Aligning Pretraining for Detection via Object-Level Contrastive LearningFangyun Wei, Yue Gao, Zhirong Wu, Han Hu et al.NeurIPS 2021 · 180 citations
