DictFormer: Tiny Transformer with Shared Dictionary
Qian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin
摘要
We introduce DictFormer with the efficient shared dictionary to provide a compact, fast, and accurate transformer model. DictFormer significantly reduces the redundancy in the transformer's parameters by replacing the prior transformer's parameters with a compact, shared dictionary, few unshared coefficients, and indices. Also, DictFormer enables faster computations since expensive weights multiplications are converted into cheap shared look-ups on dictionary and few linear projections. Training dictionary and coefficients are not trivial since indices used for looking up dictionary are not differentiable. We adopt a sparse-constraint training with relaxation to learn coefficients and indices in DictFormer. DictFormer is flexible to support different model sizes by dynamically changing dictionary size. Compared to existing lightweight Transformers, DictFormer consistently reduces model size over Transformer on multiple tasks, e.g., machine translation, abstractive summarization, and language modeling. Extensive experiments show that DictFormer reduces to model size with similar accuracy over multiple tasks, compared to Transformer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- TrojText: Test-time Invisible Textual Trojan InsertionQian Lou, Yepeng Liu, Bo FengICLR 2023 · 被引用 5 次
- DictPFL: Efficient and Private Federated Learning on Encrypted GradientsJiaqi Xue, Mayank Kumar, Yuzhang Shang, Shangqian Gao 等NeurIPS 2025 · 被引用 4 次
- PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRASheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang 等ACL 2024
- TrojViT: Trojan Insertion in Vision TransformersMengxin Zheng, Qian Lou, Lei JiangCVPR 2023
- MoS: Unleashing Parameter Efficiency of Low-Rank Adaptation with Mixture of ShardsSheng Wang, Liheng Chen, Pengan Chen, Jingwei Dong 等ICLR 2025
相关 Paper
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationSunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv 等EMNLP 2022 · 被引用 3 次
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser 等NeurIPS 2021 · 被引用 127 次
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary LearningMagauiya Zhussip, Dmitriy Shopkhoev, Ammar Ali, Stamatios LefkimmiatisAAAI 2026 · 被引用 5 次
- Sparsifying Transformer Models with Trainable Representation PoolingMichal Pietruszka, Lukasz Borchmann, Lukasz GarncarekACL 2022 · 被引用 13 次
