Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
Andres Potapczynski, Shikai Qiu, Marc Finzi, Christopher Ferri, Charlie Chen, Micah Goldblum, C. Bayan Bruss, Christopher De Sa, Andrew Gordon Wilson
摘要
Dense linear layers are the dominant computational bottleneck in large neural networks, presenting a critical need for more efficient alternatives. Previous efforts focused on a small number of hand-crafted structured matrices and neglected to investigate whether these structures can surpass dense layers in terms of compute-optimal scaling laws when both the model size and training examples are optimally allocated. In this work, we present a unifying framework that enables searching among all linear operators expressible via an Einstein summation. This framework encompasses many previously proposed structures, such as low-rank, Kronecker, Tensor-Train, Block Tensor-Train (BTT), and Monarch, along with many novel structures. To analyze the framework, we develop a taxonomy of all such operators based on their computational and algebraic properties and show that differences in the compute-optimal scaling laws are mostly governed by a small number of variables that we introduce. Namely, a small (which measures parameter sharing) and large (which measures the rank) reliably led to better scaling laws. Guided by the insight that full-rank structures that maximize parameters per unit of compute perform the best, we propose BTT-MoE, a novel Mixture-of-Experts (MoE) architecture obtained by sparsifying computation in the BTT structure. In contrast to the standard sparse MoE for each entire feed-forward network, BTT-MoE learns an MoE in every single linear layer of the model, including the projection matrices in the attention blocks. We find BTT-MoE provides a substantial compute-efficiency gain over dense layers and standard MoE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFTDa Chang, Peng Xue, Yu Li, Yongxiang Liu 等AAAI 2026 · 被引用 2 次
- EUGens: Efficient, Unified and General Dense LayersSang Min Kim, Byeongchan Kim, Arijit Sehanobish, Somnath Basu Roy Chowdhury 等NeurIPS 2025 · 被引用 1 次
- Customizing the Inductive Biases of Softmax Attention using Structured MatricesYilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu 等ICML 2025
- Scaling Probabilistic Circuits via Monarch MatricesHonghua Zhang, Meihua Dang, Benjie Wang, Stefano Ermon 等ICML 2025
- Understanding and Relaxing the Limitations of Transformers for Linear AlgebraAndres Potapczynski, Alex Ali, Andrew Gordon WilsonICLR 2026
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 被引用 214 次
- Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureDaniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson 等NeurIPS 2023 · 被引用 80 次
- The Deep Bootstrap Framework: Good Online Learners are Good Offline GeneralizersPreetum Nakkiran, Behnam Neyshabur, Hanie SedghiICLR 2021 · 被引用 75 次
- Compute Better Spent: Replacing Dense Layers with Structured MatricesShikai Qiu, Andres Potapczynski, Marc Anton Finzi, Micah Goldblum 等ICML 2024 · 被引用 26 次
相关 Paper
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsJiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang 等PPoPP 2022 · 被引用 97 次
- MoEUT: Mixture-of-Experts Universal TransformersRóbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts 等NeurIPS 2024 · 被引用 65 次
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin 等NeurIPS 2025 · 被引用 6 次
- TT-LoRA MoE: Using Parameter-Efficient Fine-Tuning and Sparse Mixture-Of-ExpertsPradip Kunwar, Minh N. Vu, Maanak Gupta, Mahmoud Abdelsalam 等SC 2025 · 被引用 1 次
- Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language ModelZeyu Liu, Tim Dettmers, Xi Lin, Veselin Stoyanov 等EMNLP 2023 · 被引用 3 次
