AlignDet: Aligning Pre-training and Fine-tuning in Object Detection
Ming Li, Jie Wu, Xionghui Wang, Chen Chen, Jie Qin, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan
摘要
The paradigm of large-scale pre-training followed by downstream fine-tuning has been widely employed in various object detection algorithms. In this paper, we reveal discrepancies in data, model, and task between the pre-training and fine-tuning procedure in existing practices, which implicitly limit the detector’s performance, generalization ability, and convergence speed. To this end, we propose AlignDet, a unified pre-training framework that can be adapted to various existing detectors to alleviate the discrepancies. AlignDet decouples the pre-training process into two stages, i.e., image-domain and box-domain pre-training. The image-domain pre-training optimizes the detection backbone to capture holistic visual abstraction, and box-domain pre-training learns instance-level semantics and task-aware concepts to initialize the parts out of the backbone. By incorporating the self-supervised pretrained backbones, we can pre-train all modules for various detectors in an unsupervised paradigm. As depicted in Figure 1, extensive experiments demonstrate that AlignDet can achieve significant improvements across diverse protocols, such as , and . For example, AlignDet improves FCOS by 5.3 mAP, RetinaNet by 2.1 mAP, Faster R-CNN by 3.3 mAP, and DETR by 2.3 mAP under fewer epochs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual ReconstructionChen-Long Duan, Yong Li, Xiu-Shen Wei, Lin ZhaoNeurIPS 2024 · 被引用 9 次
- Bi-Level Optimization for Self-Supervised AI-Generated Face DetectionMian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan 等ICCV 2025 · 被引用 2 次
- DisCo DETR: Distance-aware Multi-view Contrastive Learning for DETR Pre-trainingChao Ouyang, Yuyang Bai, Jun Zhang, Tianlu Gao 等AAAI 2026
它引用的顶会 Paper30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
相关 Paper
- DETReg: Unsupervised Pretraining with Region Priors for Object DetectionAmir Bar, Xin Wang, Vadim Kantorov, Colorado J. Reed 等CVPR 2022 · 被引用 130 次
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningYizhou Wang, Meilin Chen, Shixiang Tang, Feng Zhu 等NeurIPS 2022 · 被引用 2 次
- PreDet: Large-scale weakly supervised pre-training for detectionVignesh Ramanathan, Rui Wang, Dhruv MahajanICCV 2021 · 被引用 14 次
- Aligning Pretraining for Detection via Object-Level Contrastive LearningFangyun Wei, Yue Gao, Zhirong Wu, Han Hu 等NeurIPS 2021 · 被引用 180 次
