Deeply Tensor Compressed Transformers for End-to-End Object Detection
Peining Zhen, Ziyang Gao, Tianshu Hou, Yuan Cheng, Hai-Bao Chen
Abstract
DEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the DETR models computation and storage intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of the tensor-based compression is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we propose a gated multi-head attention (GMHA) module to mitigate the accuracy drop of the tensor-compressed DETR models. In GMHA, each attention head has an independent gate to determine the passed attention value. The redundant attention information can be suppressed by adopting the normalized gates. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO dataset to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results show that we can attain 3.7 times full model compression with 482 times feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e26fd096-90c4-484a-b1c7-7a38d0eecad1Cited by top-tier papers2
- Data-free Neural Representation Compression with Riemannian Neural DynamicsZhengqi Pei, Anran Zhang, Shuhui Wang, Xiangyang Ji et al.ICML 2024 · 5 citations
- TD-MoE: Tensor Decomposition for MoE ModelsYuebin XU, YANHONG WANG, Xuemei Peng, Hui Zang et al.ICLR 2026
Builds on9
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object DetectionYuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang et al.NeurIPS 2021 · 430 citations
- Rethinking Transformer-based Set Prediction for Object DetectionZhiqing Sun, Shengcao Cao, Yiming Yang, Kris KitaniICCV 2021 · 381 citations
Related papers
- Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRFeng Li, Ailing Zeng, Shilong Liu, Hao Zhang et al.CVPR 2023
- AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary QueriesRunqi Wang, Huixin Sun, Linlin Yang, Shaohui Lin et al.AAAI 2024 · 10 citations
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv et al.EMNLP 2022 · 3 citations
- WB-DETR: Transformer-Based Detector without BackboneFanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li et al.ICCV 2021 · 45 citations
