Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation
Sunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv, Benyou Wang, Victor Junqiu Wei, Xin Jiang
Abstract
Transformer has been demonstrated effective in Neural Machine Translation (NMT). However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback. To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, which retains full rank and meanwhile reduces operations and parameters. A Transformer using HTT, named Hypoformer, consistently and notably outperforms the recent lightweight SOTA methods on three standard translation tasks under different parameter and speed scales. In extreme low resource scenarios, Hypoformer has a 7.1 point absolute improvement in BLEU and 1.27× speedup than the vanilla Transformer on the IWSLT'14 De-En task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6bd0b84-8a5e-4c7a-b28c-d1e169c310a0Cited by top-tier papers2
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 7 citations
- MTNL: A Unified Modeling Perspective for Enhancing Tensor Network LearningJunhua Zeng, Yuning Qiu, Binghua Li, Chao Li et al.ICML 2026
Builds on7
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen et al.EMNLP 2020 · 158 citations
- DeLighT: Deep and Light-weight TransformerSachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer et al.ICLR 2021 · 96 citations
Related papers
- Deeply Tensor Compressed Transformers for End-to-End Object DetectionPeining Zhen, Ziyang Gao, Tianshu Hou, Yuan Cheng et al.AAAI 2022 · 19 citations
- DictFormer: Tiny Transformer with Shared DictionaryQian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen et al.ICLR 2022 · 12 citations
- Learning Light-Weight Translation Models from Deep TransformerBei Li, Ziyang Wang, Hui Liu, Quan Du et al.AAAI 2021 · 44 citations
- Towards Lightweight Time Series Forecasting: A Patch-Wise Transformer with Weak Data EnrichingMeng Wang, Jintao Yang, Bin Yang, Hui Li et al.ICDE 2025 · 10 citations
- Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsSean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto et al.SIGIR 2020 · 62 citations
