Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache Compression
Xuefei Wang, Haoyu Tang, Tianyuan Liang, Zhibin Wang, Yupeng Hu, Weili Guan
摘要
The linear growth of KV cache bottlenecks long-context LLMs, yet RoPE-induced oscillations complicate Key cache quantization. To address this issue, we propose SpectrumQuant, a frequency-domain framework that utilizes the Discrete Cosine Transform (DCT) to convert these oscillations into sparse spectral representations. Specifically, our pipeline integrates dominant frequency extraction, hybrid bit-width allocation, and high-frequency preemphasis to maximize fidelity while minimizing memory footprint. To eliminate computational overhead, we develop fused Triton kernels featuring deferred inverse transformation and on-chip sparse accumulation. Extensive experiments on several benchmarks confirm SpectrumQuant achieves efficient compression with performance and latency comparable to FP16 baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney 等NeurIPS 2024 · 被引用 738 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank ControlPriyansh Bhatnagar, Ashkan Moradifirouzabadi, Se-Hyun Yang, SeungJae Lee 等ICML 2026 · 被引用 3 次
- CommVQ: Commutative Vector Quantization for KV Cache CompressionJunyan Li, Yang Zhang, Muhammad Yusuf Hassan, Talha Chafekar 等ICML 2025
- SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs QuantizationZhixiong Zhao, Fangxin Liu, Junjie Wang, Chenyang Guan 等AAAI 2026 · 被引用 4 次
- JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context InferenceChengyu Sun, Yaqi Xia, Hulin Wang, Donglin Yang 等PPoPP 2026 · 被引用 1 次
- SALS: Sparse Attention in Latent Space for KV Cache CompressionJunlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu 等NeurIPS 2025 · 被引用 7 次
