WB-DETR: Transformer-Based Detector without Backbone
Fanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li, Jingquan Peng, Zihao Li
摘要
Transformer-based detector is a new paradigm in object detection, which aims to achieve pretty-well performance while eliminates the priori knowledge driven components, e.g., anchors, proposals and the NMS. DETR, the state-of-the-art model among them, is composed of three sub-modules, i.e., a CNN-based backbone and paired transformer encoder-decoder. The CNN is applied to extract local features and the transformer is used to capture global contexts. This pipeline, however, is not concise enough. In this paper, we propose WB-DETR (DETR-based detector Without Backbone) to prove that the reliance on CNN features extraction for a transformer-based detector is not necessary. Unlike the original DETR, WB-DETR is composed of only an encoder and a decoder without CNN backbone. For an input image, WB-DETR serializes it directly to encode the local features into each individual token. To make up the deficiency of transformer in modeling local information, we design an LIE-T2T (local information enhancement tokens to token) module to enhance the internal information of tokens after unfolding. Experimental results demonstrate that WB-DETR, the first pure-transformer detector without CNN to our knowledge, yields on par accuracy and faster inference speed with only half number of parameters compared with DETR baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Cascade-DETR: Delving into High-Quality Universal Object DetectionMingqiao Ye, Lei Ke, Siyuan Li, Yu-Wing Tai 等ICCV 2023 · 被引用 62 次
- Sparse Semi-DETR: Sparse Learnable Queries for Semi-Supervised Object DetectionTahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan AfzalCVPR 2024 · 被引用 36 次
- SpeedDETR: Speed-aware Transformers for End-to-end Object DetectionPeiyan Dong, Zhenglun Kong, Xin Meng, Peng Zhang 等ICML 2023 · 被引用 21 次
- R(Det)2: Randomized Decision Routing for Object DetectionYali Li, Shengjin WangCVPR 2022 · 被引用 9 次
- 1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense PredictionsDongshuo Yin, Yiran Yang, Zhechao Wang, Hongfeng Yu 等CVPR 2023
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
相关 Paper
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang 等CVPR 2022 · 被引用 14 次
- DECO: Unleashing the Potential of ConvNets for Query-based Detection and SegmentationXinghao Chen, Siwei Li, Yijing Yang, Yunhe WangICLR 2025
- DA-DETR: Domain Adaptive Detection Transformer with Information FusionJingyi Zhang, Jiaxing Huang, Zhipeng Luo, Gongjie Zhang 等CVPR 2023
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
