EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq Generation
Tao Ge, Si-Qing Chen, Furu Wei
Abstract
We introduce EdgeFormer – a parameter-efficient Transformer for on-device seq2seq generation under the strict computation and memory constraints. Compared with the previous parameter-efficient Transformers, EdgeFormer applies two novel principles for cost-effective parameterization, allowing it to perform better given the same parameter budget; moreover, EdgeFormer is further enhanced by layer adaptation innovation that is proposed for improving the network with shared layers.Extensive experiments show EdgeFormer can effectively outperform previous parameter-efficient Transformer baselines and achieve competitive results under both the computation and memory constraints. Given the promising results, we release EdgeLM – the pretrained version of EdgeFormer, which is the first publicly available pretrained on-device seq2seq model that can be easily fine-tuned for seq2seq tasks with strong results, facilitating on-device seq2seq generation in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ed576a6-cfb1-483a-b6ed-7e87e40d9424Cited by top-tier papers5
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
- PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRASheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang et al.ACL 2024
- Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRASangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji et al.ICLR 2025
- MoS: Unleashing Parameter Efficiency of Low-Rank Adaptation with Mixture of ShardsSheng Wang, Liheng Chen, Pengan Chen, Jingwei Dong et al.ICLR 2025
- GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on GraphsYi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding et al.CVPR 2025
Builds on13
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 167 citations
Related papers
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer VotingZhongzhi Yu, Zheng Wang, Yuhan Li, Ruijie Gao et al.DAC 2024 · 57 citations
- WeGeFT: Weight‑Generative Fine-Tuning for Multi-Faceted Efficient Adaptation of Large ModelsChinmay Savadikar, Xi Song, Tianfu WuICML 2025
- Towards Lightweight Time Series Forecasting: A Patch-Wise Transformer with Weak Data EnrichingMeng Wang, Jintao Yang, Bin Yang, Hui Li et al.ICDE 2025 · 10 citations
- LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory UnitsZeyu Liu, Gourav Datta, Anni Li, Peter Anthony BeerelICLR 2024 · 18 citations
- AdMiT: Adaptive Multi-Source Tuning in Dynamic EnvironmentsXiangyu Chang, Fahim Faisal Niloy, Sk Miraj Ahmed, Srikanth V. Krishnamurthy et al.CVPR 2025
