Lune

ISCA2026顶会

Approaching Shannon Bound with Lossless LLM Weight Compression

Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He

2026年份
2被引次数

摘要

Large language models (LLMs) now scale to trillions of parameters, driving weight storage into the terabyte regime and creating an acute mismatch with GPU memory capacity. Although lossless compression is widely effective in other domains, it remains underutilized in LLM systems. Through a comprehensive entropy study across models from 1.5B to 405B parameters and numeric formats ranging from bf16 to int4 and AWQ/SQ8, we find that LLM weights contain far less intrinsic randomness than their stored bitwidth implies, their effective entropy is 2-10× lower, indicating that up to a 10×\mathbf{1 0} \times footprint reduction is theoretically achievable without altering any weight values. Leveraging this insight, we introduce a tile-level, on-thefly lossless decompression framework based on Asymmetric Numeral Systems that aligns decoding with the GEMM tiling pattern of GPU inference. Our design achieves bit-rates within 0.01-0.1 bits of the Shannon limit across a wide range of LLM numerical formats, demonstrating that nearly all statistical redundancy is eliminated. Integrated into the SGLang serving framework with multi-GPU support, our approach increases the maximum batch size of Qwen-14B from 47 to 75, improving throughput by up to 1.2×\mathbf{1. 2} \times. On Mixtral-176B, the feasible batch size increases from 20 to 95 (4.8×), yielding up to 1.6×\mathbf{1. 6} \times throughput improvement. Compared to state-of-the-art lossless compression approaches NeuZip and DFloat11, our design further improves throughput by up to 11×\mathbf{1 1} \times.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f39e9083-0e2d-41e5-9f6f-f02d5d402ce6

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖