Lune

PPoPP2026Top-tier venue

JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context Inference

Chengyu Sun, Yaqi Xia, Hulin Wang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng

2026Year
1Citations

Abstract

Long-context large language models (LLMs) have seen widespread adoption in recent years. However, during inference, the key-value (KV) cache—which stores intermediate activations—consumes significant memory, particularly as sequence lengths grow. Quantization offers a promising path to compress KV cache, but existing 2-bit approaches fall short of achieving optimal inference efficiency due to hardware-unfriendly algorithms and system implementations.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get fa2d59cf-413f-464f-9898-edf95506e45b

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines