Lizard: An Efficient Linearization Framework for Large Language Models
Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen
摘要
We propose Lizard, a linearization framework that transforms pretrained Transformer-based Large Language Models (LLMs) into subquadratic architectures. Transformers faces severe computational and memory bottlenecks with long sequences due to the quadratic complexity of softmax attention and the growing Key-Value (KV) cache that makes inference memory-bound by context length. Lizard addresses these limitations by introducing a subquadratic attention mechanism that closely approximates softmax attention while preserving model quality. Unlike prior linearization methods constrained by fixed, non-adaptive structures, Lizard augments the architecture with compact, learnable modules that enable adaptive memory control and robust length generalization. Moreover, we introduce a hardwareaware algorithm that solves numerical instability in gated attention to accelerate training. Extensive experiments show that Lizard achieves near-lossless recovery of its teacher model's performance, significantly outperforming previous methods by up to 9.4 - 24.5 points on the 5-shot MMLU benchmark and demonstrating superior associative recall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Distilling to Hybrid Attention Models via KL-Guided Layer SelectionYanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra 等ICLR 2026 · 被引用 17 次
- Effective Distillation to Hybrid xLSTM ArchitecturesLukas Hauzenberger, Niklas Schmidinger, Thomas Schmied, Anamaria-Roberta Hartl 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Liger: Linearizing Large Language Models to Gated Recurrent StructuresDisen Lan, Weigao Sun, Jiaxi Hu, Jusen Du 等ICML 2025
- LoLCATs: On Low-Rank Linearizing of Large Language ModelsMichael Zhang, Simran Arora, Rahul Chalamala, Benjamin Frederick Spector 等ICLR 2025
- PolySketchFormer: Fast Transformers via Sketching Polynomial KernelsPraneeth Kacham, Vahab Mirrokni, Peilin ZhongICML 2024 · 被引用 27 次
- The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax MimicryMichael Zhang, Kush Bhatia, Hermann Kumbong, Christopher RéICLR 2024 · 被引用 103 次
- Luna: Linear Unified Nested AttentionXuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou 等NeurIPS 2021 · 被引用 145 次
