FrameQuant: Flexible Low-Bit Quantization for Transformers
Harshavardhan Adepu, Zhanpeng Zeng, Li Zhang, Vikas Singh
Abstract
Transformers are the backbone of powerful foundation models for many Vision and Natural Language Processing tasks. But their compute and memory/storage footprint is large, and so, serving such models is expensive often requiring high-end hardware. To mitigate this difficulty, Post-Training Quantization seeks to modify a pre-trained model and quantize it to eight bits or lower, significantly boosting compute/memory/latency efficiency. Such models have been successfully quantized to four bits with some performance loss. In this work, we outline a simple scheme to quantize Transformer-based models to just two bits (plus some overhead) with only a small drop in accuracy. Key to our formulation is a concept borrowed from Harmonic analysis called Fusion Frames. Our main finding is that the quantization must take place not in the original weight space, but instead in the Fusion Frame representations. If quantization is interpreted as the addition of noise, our casting of the problem allows invoking an extensive body of known consistent recovery and noise robustness guarantees. Further, if desired, de-noising filters are known in closed form. We show empirically, via a variety of experiments, that (almost) two-bit quantization for Transformer models promises sizable efficiency gains. The code is available at https://github.com/vsingh-group/FrameQuant
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ad58bd5-2e6a-4620-98c4-caeedfc11ef3Cited by top-tier papers6
- MagR: Weight Magnitude Reduction for Enhancing Post-Training QuantizationAozhong Zhang, Naigang Wang, Yanxia Deng, Xin Li et al.NeurIPS 2024 · 33 citations
- Qronos: Correcting the Past by Shaping the Future... in Post-Training QuantizationShihao Zhang, Haoyu Zhang, Ian Colbert, Rayan SaabICLR 2026 · 27 citations
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in InferenceHao Zhang, Mengsi Lyu, Zhuo Chen, Yulong Ao et al.ACL 2026 · 9 citations
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMsNingning Chen, Weicai Ye, Ying JiangNeurIPS 2025 · 5 citations
- Matryoshka QuantizationPranav Ajit Nair, Puranjay Datta, Jeff Dean, Prateek Jain et al.ICML 2025
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Mr.BiQ: Post-Training Non-Uniform Quantization based on Minimizing the Reconstruction ErrorYongkweon Jeon, Chungman Lee, Eulrang Cho, Yeonju RoCVPR 2022 · 28 citations
- 8-bit Transformer Inference and Fine-tuning for Edge AcceleratorsJeffrey Yu, Kartik Prabhu, Yonatan Urman, Robert M. Radway et al.ASPLOS 2024 · 26 citations
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 10 citations
- Towards Accurate Post-Training Quantization for Vision TransformerYifu Ding, Haotong Qin, Qinghua Yan, Zhenhua Chai et al.ACM MM 2022 · 68 citations
- Understanding and Overcoming the Challenges of Efficient Transformer QuantizationYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortEMNLP 2021 · 74 citations
