SingularBit: Exploiting Synergy of Singular Value Decomposition and Low-Bit Quantization for Weight and KV Compression in LLM Inference
Seongyon Hong, Hyundeok Kong, Jungwan Lee, Sangjin Kim, Hoi-Jun Yoo
摘要
Large language models (LLMs) have become indispensable across diverse domains, while inference-time scaling has further enhanced their ability to tackle complex real-world tasks through long-context processing and extended generation. However, autoregressive decoding with extended sequences intensifies the memory bandwidth bottleneck through repeated weight access and growing key-value (KV) caches, a challenge we term the Dual Memory Wall. Although quantization and low-rank approximation have been explored independently for weight and KV compression, existing approaches remain insufficient for the demands of modern LLM inference. We present SingularBit, the first accelerator to exploit the synergy between singular value decomposition (SVD) and low-bit quantization for both offline weight and online KV cache compression. Our key insight is that singular values decay rapidly in both weight matrices and attention distributions, enabling aggressive compression of less critical components while preserving model accuracy through rank-aware mixed-precision allocation. SingularBit comprises two algorithm-hardware co-designs. First, SingularBit-W applies rank-aware mixed-precision quantization to weights, compressing most parameters to 1-2 bits guided by singular value magnitude, executed efficiently by the SingularBit Tensor Core supporting rank-wise mixed-precision computation. Second, SingularBit-KV extends this principle to online KV cache compression, dynamically allocating precision based on both token importance and rank significance, accelerated by the SingularBit Compression Engine. Synthesized in 28nm CMOS, SingularBit achieves stateof-the-art algorithmic efficiency across standard language modeling, commonsense reasoning, and long-context benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV CacheFei Li, Song Liu, Weiguo Wu, Shiqiang Nie 等AAAI 2026 · 被引用 1 次
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 被引用 14 次
- STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank ControlPriyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang, SeungJae Lee 等ICML 2026 · 被引用 3 次
- TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM QuantizationHaodong WANG, Junjie Liu, Zicong Hong, Qianli Liu 等ICML 2026
- MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context ReasoningTao Zhang, Ziqian Zeng, Hao Peng, Huiping Zhuang 等ACL 2026 · 被引用 3 次
