SingularBit: Exploiting Synergy of Singular Value Decomposition and Low-Bit Quantization for Weight and KV Compression in LLM Inference
Seongyon Hong, Hyundeok Kong, Jungwan Lee, Sangjin Kim, Hoi-Jun Yoo
Abstract
Large language models (LLMs) have become indispensable across diverse domains, while inference-time scaling has further enhanced their ability to tackle complex real-world tasks through long-context processing and extended generation. However, autoregressive decoding with extended sequences intensifies the memory bandwidth bottleneck through repeated weight access and growing key-value (KV) caches, a challenge we term the Dual Memory Wall. Although quantization and low-rank approximation have been explored independently for weight and KV compression, existing approaches remain insufficient for the demands of modern LLM inference. We present SingularBit, the first accelerator to exploit the synergy between singular value decomposition (SVD) and low-bit quantization for both offline weight and online KV cache compression. Our key insight is that singular values decay rapidly in both weight matrices and attention distributions, enabling aggressive compression of less critical components while preserving model accuracy through rank-aware mixed-precision allocation. SingularBit comprises two algorithm-hardware co-designs. First, SingularBit-W applies rank-aware mixed-precision quantization to weights, compressing most parameters to 1-2 bits guided by singular value magnitude, executed efficiently by the SingularBit Tensor Core supporting rank-wise mixed-precision computation. Second, SingularBit-KV extends this principle to online KV cache compression, dynamically allocating precision based on both token importance and rank significance, accelerated by the SingularBit Compression Engine. Synthesized in 28nm CMOS, SingularBit achieves stateof-the-art algorithmic efficiency across standard language modeling, commonsense reasoning, and long-context benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9bd5bb45-35d0-41e3-85a1-a40eea224eb3Related papers
- KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV CacheFei Li, Song Liu, Weiguo Wu, Shiqiang Nie et al.AAAI 2026 · 1 citation
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank ControlPriyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang, SeungJae Lee et al.ICML 2026 · 3 citations
- TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM QuantizationHaodong WANG, Junjie Liu, Zicong Hong, Qianli Liu et al.ICML 2026
- MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context ReasoningTao Zhang, Ziqian Zeng, Hao Peng, Huiping Zhuang et al.ACL 2026 · 3 citations
