RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision Quantization
Qihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei, Jidong Zhai
2026Year
Abstract
Mixed precision quantization has been adopted to accelerate large language models (LLMs) serving by leveraging high-throughput low-precision compute units in GPUs while preserving outliers in higher precision to maintain model accuracy. However, existing methods focus on mitigating single-dimensional channel-wise outliers, leading to model accuracy degradation when scaled to 4-bit precision.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c57fd1c2-70ab-4760-8af3-8cd52ab035e6Related papers
- OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language ModelsChanghun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim et al.AAAI 2024 · 134 citations
- Oltron: Algorithm-Hardware Co-design for Outlier-Aware Quantization of LLMs with Inter-/Intra-Layer AdaptationChenhao Xue, Chen Zhang, Xun Jiang, Zhutianya Gao et al.DAC 2024 · 11 citations
- Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and VariationsPatrick Blumenberg, Thomas Graave, Tim FingscheidtICLR 2026 · 5 citations
- Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language ModelsJung Hwan Heo, Jeonghoon Kim, Beomseok Kwon, Byeongwook Kim et al.ICLR 2024
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu et al.ASPLOS 2025 · 5 citations
