Rethinking Transformer-based Set Prediction for Object Detection
Zhiqing Sun, Shengcao Cao, Yiming Yang, Kris Kitani
Abstract
DETR is a recently proposed Transformer-based method which views object detection as a set prediction problem and achieves state-of-the-art performance but demands extra-long training time to converge. In this paper, we investigate the causes of the optimization difficulty in the training of DETR. Our examinations reveal several factors contributing to the slow convergence of DETR, primarily the issues with the Hungarian loss and the Transformer cross-attention mechanism. To overcome these issues we propose two solutions, namely, TSP-FCOS (Transformer-based Set Prediction with FCOS) and TSP-RCNN (Transformer-based Set Prediction with RCNN). Experimental results show that the proposed methods not only converge much faster than the original DETR, but also significantly outperform DETR and other baselines in terms of detection accuracy. Code is released at https://github.com/Edward-Sun/TSP-Detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82b804d6-fda4-4de8-8ac4-10cc5faf11efCited by top-tier papers69
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Builds on12
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- Scale-Aware Trident Networks for Object DetectionYanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 1,031 citations
Related papers
- Accelerating DETR Convergence via Semantic-Aligned MatchingGongjie Zhang, Zhipeng Luo, Yingchen Yu, Kaiwen Cui et al.CVPR 2022 · 116 citations
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang et al.ICCV 2023 · 231 citations
- Fast Convergence of DETR with Spatially Modulated Co-AttentionPeng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai et al.ICCV 2021 · 392 citations
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 59 citations
- DAC-DETR: Divide the Attention Layers and ConquerZhengdong Hu, Yifan Sun, Jingdong Wang, Yi YangNeurIPS 2023 · 53 citations
