An Efficient Transformer Decoder with Compressed Sub-layers
Yanyang Li, Ye Lin, Tong Xiao, Jingbo Zhu
Abstract
The large attention-based encoder-decoder network (Transformer) has become prevailing recently due to its effectiveness. But the high computation complexity of its decoder raises the inefficiency issue. By examining the mathematic formulation of the decoder, we show that under some mild conditions, the architecture could be simplified by compressing its sub-layers, the basic building block of Transformer, and achieves a higher parallelism. We thereby propose Compressed Attention Network, whose decoder layer consists of only one sub-layer instead of three. Extensive experiments on 14 WMT machine translation tasks show that our model is 1.42x faster with performance on par with a strong baseline. This strong baseline is already 2x faster than the widely used standard baseline without loss in performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Towards Continual Knowledge Learning of Language ModelsJoel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin et al.ICLR 2022 · 204 citations
- UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data ScienceYazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu et al.ICLR 2024 · 35 citations
- RankNAS: Efficient Neural Architecture Search by Pairwise RankingChi Hu, Chenglong Wang, Xiangnan Ma, Xia Meng et al.EMNLP 2021 · 10 citations
- Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyYanyang Li, Fuli Luo, Runxin Xu, Songfang Huang et al.ACL 2022 · 3 citations
- Efficient Inference for Multilingual Neural Machine TranslationAlexandre Berard, Dain Lee, Stéphane Clinchant, Kweon Woo Jung et al.EMNLP 2021
Builds on5
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Shallow-to-Deep Training for Neural Machine TranslationBei Li, Ziyang Wang, Hui Liu, Yufan Jiang et al.EMNLP 2020 · 41 citations
- Neural Machine Translation with Joint RepresentationYanyang Li, Qiang Wang, Tong Xiao, Tongran Liu et al.AAAI 2020 · 10 citations
Related papers
- Glancing Transformer for Non-Autoregressive Neural Machine TranslationLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang et al.ACL 2021
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 52 citations
- Attention-Only Transformers via Unrolled Subspace DenoisingPeng Wang, Yifu Lu, Yaodong Yu, Druv Pai et al.ICML 2025
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser et al.NeurIPS 2021 · 127 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
