Irrational Complex Rotations Empower Low-bit Optimizers
Zhen Tian, Xin Zhao, Ji-Rong Wen
摘要
In this paper, we propose a novel optimizer state compression algorithm, namely -Quant, which leverages the properties of irrational numbers (e.g., ) for memory-efficient training. The core idea is based on our mathematical findings, which show that a pair of parameters can be represented by a single rotation angle using the complex rotation scheme. Building on this insight, we map the parameters into a complex space and perform quantization using the corresponding rotation angles. To efficiently integrate it into optimization process, we develop an efficient system of geometric equations that computes the precise rotation angles with linear complexity. We evaluate -Quant on a wide range of tasks. Our experiments show that it can reduce the bit-width of parameters to 3.32-bit, achieving a 75% reduction in parameter scale and a 40% decrease in GPU memory usage, all while maintaining full accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- An Image Patch is a Wave: Phase-Aware Vision MLPYehui Tang, Kai Han, Jianyuan Guo, Chang Xu 等CVPR 2022 · 被引用 137 次
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 被引用 91 次
相关 Paper
- Memory Efficient Optimizers with 4-bit StatesBingrui Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 72 次
- Achieving low-bit Muon through subspace preservation and grid quantizationHuaijin Wu, Bingrui Li, Yebin Yang, Yi Tu 等ICLR 2026
- FlashOptim: Memory Efficient Optimizers for Large-Scale TrainingJose Javier Gonzalez Ortiz, Abhay Gupta, Christopher Rinard, Davis BlalockICML 2026
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 被引用 19 次
- 4-bit Shampoo for Memory-Efficient Network TrainingSike Wang, Pan Zhou, Jia Li, Hua HuangNeurIPS 2024 · 被引用 19 次
