Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions
Stefano Massaroli, Michael Poli, Daniel Y. Fu, Hermann Kumbong, Rom N. Parnichkun, David W. Romero, Aman Timalsina, Quinn McIntyre, Beidi Chen, Atri Rudra, Ce Zhang, Christopher Ré
Abstract
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads -- naively requiring a full pass (or caching of activations) over the input sequence for each generated token -- similarly to attention-based models. In this paper, we seek to enable compute and memory cost per token in any pre-trained long convolution architecture to reduce memory footprint and increase throughput during generation. Concretely, our methods consist in extracting low-dimensional linear state-space models from each convolution layer, building upon rational interpolation and model-order reduction techniques. We further introduce architectural improvements to convolution-based layers such as Hyena: by weight-tying the filters across channels into heads, we achieve higher pre-training quality and reduce the number of filters to be distilled. The resulting model achieves 10x higher throughput than Transformers and 1.5x higher than Hyena at 1.3B parameters, without any loss in quality after distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- The Mamba in the Llama: Distilling and Accelerating Hybrid ModelsJunxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush et al.NeurIPS 2024 · 146 citations
- Zoology: Measuring and Improving Recall in Efficient Language ModelsSimran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson et al.ICLR 2024 · 140 citations
- Mechanistic Design and Scaling of Hybrid ArchitecturesMichael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy et al.ICML 2024 · 57 citations
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing et al.ICLR 2026 · 41 citations
Builds on15
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
Related papers
- Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and BeyondCostin-Andrei Oncescu, Sanket Purandare, Stratos Idreos, Sham M. KakadeICLR 2025
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- State-Free Inference of State-Space Models: The Transfer Function ApproachRom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith et al.ICML 2024 · 18 citations
- Geometric Hyena Networks for Large-scale Equivariant LearningArtem Moskalev, Mangal Prakash, Junjie Xu, Tianyu Cui et al.ICML 2025
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen et al.ICLR 2025
