DictFormer: Tiny Transformer with Shared Dictionary
Qian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin
Abstract
We introduce DictFormer with the efficient shared dictionary to provide a compact, fast, and accurate transformer model. DictFormer significantly reduces the redundancy in the transformer's parameters by replacing the prior transformer's parameters with a compact, shared dictionary, few unshared coefficients, and indices. Also, DictFormer enables faster computations since expensive weights multiplications are converted into cheap shared look-ups on dictionary and few linear projections. Training dictionary and coefficients are not trivial since indices used for looking up dictionary are not differentiable. We adopt a sparse-constraint training with relaxation to learn coefficients and indices in DictFormer. DictFormer is flexible to support different model sizes by dynamically changing dictionary size. Compared to existing lightweight Transformers, DictFormer consistently reduces model size over Transformer on multiple tasks, e.g., machine translation, abstractive summarization, and language modeling. Extensive experiments show that DictFormer reduces to model size with similar accuracy over multiple tasks, compared to Transformer.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d5279ee5-8810-4a47-8fef-7b593b04535dCited by top-tier papers6
- TrojText: Test-time Invisible Textual Trojan InsertionQian Lou, Yepeng Liu, Bo FengICLR 2023 · 5 citations
- DictPFL: Efficient and Private Federated Learning on Encrypted GradientsJiaqi Xue, Mayank Kumar, Yuzhang Shang, Shangqian Gao et al.NeurIPS 2025 · 4 citations
- PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRASheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang et al.ACL 2024
- TrojViT: Trojan Insertion in Vision TransformersMengxin Zheng, Qian Lou, Lei JiangCVPR 2023
- MoS: Unleashing Parameter Efficiency of Low-Rank Adaptation with Mixture of ShardsSheng Wang, Liheng Chen, Pengan Chen, Jingwei Dong et al.ICLR 2025
Related papers
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 60 citations
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv et al.EMNLP 2022 · 3 citations
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser et al.NeurIPS 2021 · 127 citations
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary LearningMagauiya Zhussip, Dmitriy Shopkhoev, Ammar Ali, Stamatios LefkimmiatisAAAI 2026 · 5 citations
- Sparsifying Transformer Models with Trainable Representation PoolingMichal Pietruszka, Lukasz Borchmann, Lukasz GarncarekACL 2022 · 13 citations
