Block-Recurrent Transformers
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, Behnam Neyshabur
Abstract
We introduce the Block-Recurrent Transformer, which applies a transformer layer in a recurrent fashion along a sequence, and has linear complexity with respect to sequence length. Our recurrent cell operates on blocks of tokens rather than single tokens during training, and leverages parallel computation within a block in order to make efficient use of accelerator hardware. The cell itself is strikingly simple. It is merely a transformer layer: it uses self-attention and cross-attention to efficiently compute a recurrent function over a large set of state vectors and tokens. Our design was inspired in part by LSTM cells, and it uses LSTM-style gates, but it scales the typical LSTM cell up by several orders of magnitude. Our implementation of recurrence has the same cost in both computation time and parameter count as a conventional transformer layer, but offers dramatically improved perplexity in language modeling tasks over very long sequences. Our model out-performs a long-range Transformer XL baseline by a wide margin, while running twice as fast. We demonstrate its effectiveness on PG19 (books), arXiv papers, and GitHub source code. Our code has been released as open source [1] .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a838a96-e8c4-49d4-8fc8-5fca50ed1061Cited by top-tier papers32
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 368 citations
- Test-Time Training Done RightTianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang et al.ICLR 2026 · 127 citations
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthXuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen et al.NeurIPS 2024 · 63 citations
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon et al.NeurIPS 2024 · 52 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
Related papers
- Block-State TransformersJonathan Pilault, Mahan Fathi, Orhan Firat, Chris Pal et al.NeurIPS 2023 · 33 citations
- Blockwise Parallel Transformers for Large Context ModelsHao Liu, Pieter AbbeelNeurIPS 2023 · 55 citations
- Transformer Quality in Linear TimeWeizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. LeICML 2022 · 335 citations
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 252 citations
- Multi-Head Highly Parallelized LSTM Decoder for Neural Machine TranslationHongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong et al.ACL 2021
