Deformable Butterfly: A Highly Structured and Sparse Linear Transform
Rui Lin, Jie Ran, King Hung Chiu, Graziano Chesi, Ngai Wong
Abstract
We introduce a new kind of linear transform named Deformable Butterfly (DeBut) that generalizes the conventional butterfly matrices and can be adapted to various input-output dimensions. It inherits the fine-to-coarse-grained learnable hierarchy of traditional butterflies and when deployed to neural networks, the prominent structures and sparsity in a DeBut layer constitutes a new way for network compression. We apply DeBut as a drop-in replacement of standard fully connected and convolutional layers, and demonstrate its superiority in homogenizing a neural network and rendering it favorable properties such as light weight and low inference complexity, without compromising accuracy. The natural complexity-accuracy tradeoff arising from the myriad deformations of a DeBut layer also opens up new rooms for analytical and practical research. The codes and Appendix are publicly available at: https://github.com/ruilin0212/DeBut .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef05e384-b495-40f6-acb2-b97d195c3abcCited by top-tier papers4
- Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingTri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai et al.ICML 2022 · 125 citations
- Simple Hardware-Efficient Long Convolutions for Sequence ModelingDaniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas et al.ICML 2023 · 72 citations
- Does a sparse ReLU network training problem always admit an optimum ?Quoc-Tung Le, Rémi Gribonval, Elisa RicciettiNeurIPS 2023 · 5 citations
- Fast Inference with Kronecker-Sparse MatricesAntoine Gonon, Léon Zheng, Pascal Carrivain, Quoc-Tung LeICML 2025
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear MapsTri Dao, Nimit Sharad Sohoni, Albert Gu, Matthew Eichhorn et al.ICLR 2020 · 55 citations
Related papers
- ButterflyFlow: Building Invertible Layers with Butterfly MatricesChenlin Meng, Linqi Zhou, Kristy Choi, Tri Dao et al.ICML 2022 · 13 citations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network ModelsBeidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang et al.ICLR 2022 · 94 citations
- FSNet: Compression of Deep Convolutional Neural Networks by Filter SummaryYingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan et al.ICLR 2020 · 19 citations
- Diverse Branch Block: Building a Convolution as an Inception-Like UnitXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2021
- DeLighT: Deep and Light-weight TransformerSachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer et al.ICLR 2021 · 96 citations
