DeLighT: Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, Hannaneh Hajishirzi
Abstract
We introduce a deep and light-weight transformer, DeLighT, that delivers similar or better performance than standard transformer-based models with significantly fewer parameters. DeLighT more efficiently allocates parameters both (1) within each Transformer block using the DeLighT transformation, a deep and lightweight transformation and (2) across blocks using block-wise scaling, that allows for shallower and narrower DeLighT blocks near the input and wider and deeper DeLighT blocks near the output. Overall, DeLighT networks are 2.5 to 4 times deeper than standard transformer models and yet have fewer parameters and operations. Experiments on benchmark machine translation and language modeling tasks show that DeLighT matches or improves the performance of baseline Transformers with 2 to 3 times fewer parameters on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Filter-enhanced MLP is All You Need for Sequential RecommendationKun Zhou, Hui Yu, Wayne Xin Zhao, Ji-Rong WenWWW 2022 · 411 citations
- Rethinking Mobile Block for Efficient Attention-based ModelsJiangning Zhang, Xiangtai Li, Jian Li, Liang Liu et al.ICCV 2023 · 223 citations
- A Self-Correcting Sequential RecommenderYujie Lin, Chenyang Wang, Zhumin Chen, Zhaochun Ren et al.WWW 2023 · 31 citations
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual FlowXinlei Yu, Chengming Xu, Guibin Zhang, Yongbo He et al.ICLR 2026 · 15 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
Related papers
- DeFINE: Deep Factorized Input Token Embeddings for Neural Sequence ModelingSachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, Hannaneh HajishirziICLR 2020 · 28 citations
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du et al.AAAI 2021 · 44 citations
- Weight Distillation: Transferring the Knowledge in Neural Network ParametersYe Lin, Yanyang Li, Ziyang Wang, Bei Li et al.ACL 2021
- Go Wider Instead of DeeperFuzhao Xue, Ziji Shi, Futao Wei, Yuxuan Lou et al.AAAI 2022 · 106 citations
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
