Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation
Maximiliana Behnke, Kenneth Heafield
Abstract
The attention mechanism is the crucial component of the transformer architecture. Recent research shows that most attention heads are not confident in their decisions and can be pruned after training. However, removing them before training a model results in lower quality. In this paper, we apply the lottery ticket hypothesis to prune heads in the early stages of training, instead of doing so on a fully converged model. Our experiments on machine translation show that it is possible to remove up to three-quarters of all attention heads from a transformer-big model with an average -0.1 change in BLEU for Turkish→English. The pruned model is 1.5 times as fast at inference, albeit at the cost of longer training. The method is complementary to other approaches, such as teacher-student, with our English→German student losing 0.2 BLEU at 75% encoder attention sparsity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f37656e1-2fdc-4b6f-9d15-bc1cf1df987cCited by top-tier papers15
- Validating the Lottery Ticket Hypothesis with Inertial Manifold TheoryZeru Zhang, Jiayin Jin, Zijie Zhang, Yang Zhou et al.NeurIPS 2021 · 45 citations
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical SystemsSubhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, Tanmoy ChakrabortyNeurIPS 2021 · 34 citations
- Fast Attention Over Long Sequences With Dynamic Sparse Flash AttentionMatteo Pagliardini, Daniele Paliotta, Martin Jaggi, François FleuretNeurIPS 2023 · 26 citations
- FEASTA: A Flexible and Efficient Accelerator for Sparse Tensor Algebra in Machine LearningKai Zhong, Zhenhua Zhu, Guohao Dai, Hongyi Wang et al.ASPLOS 2024 · 16 citations
- Where does In-context Learning Happen in Large Language Models?Suzanna Sia, David Mueller, Kevin DuhNeurIPS 2024 · 14 citations
Builds on1
Related papers
- When BERT Plays the Lottery, All Tickets Are WinningSai Prasanna, Anna Rogers, Anna RumshiskyEMNLP 2020 · 114 citations
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 55 citations
- Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual TranslationBinbin Liu, Wenhan Han, Feng Chen, Yifan Zhang et al.ICLR 2026
- Sparsifying Transformer Models with Trainable Representation PoolingMichal Pietruszka, Lukasz Borchmann, Lukasz GarncarekACL 2022 · 13 citations
- Contributions of Transformer Attention Heads in Multi- and Cross-lingual TasksWeicheng Ma, Kai Zhang, Renze Lou, Lili Wang et al.ACL 2021
