Synchronizing Probabilities in Model-Driven Lossless Compression
Aviv Adler, Jennifer Tang
Abstract
It is well-known in the field of lossless data compression that probabilistic next-symbol prediction can be used to compress sequences of symbols. Deep neural networks are able to capture rich dependencies in data, offering a powerful means of estimating these probabilities and hence an avenue towards more effective compression algorithms. However, both compressor and decompressor must have exactly matching predictions; even small differences from non-determinism (which often happen with learned models due to hardware, software, or computation order) can lead to cascading decoding failures. In this paper, we formalize the problem of prediction mismatch in model-driven compression, and introduce Probability Matching Interval Coding (PMATIC), a model-agnostic algorithm that tolerates bounded prediction mismatch with low overhead. PMATIC works with the predicted probabilities, making it compatible as a drop-in replacement for the arithmetic encoder in model-driven compression tools. We show theoretical correctness and performance bounds for PMATIC, and validate these results on text data. These results confirm that, when paired an advanced prediction model, PMATIC is robust to prediction mismatch while achieving compression rates that out-perform standard modern compression tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df75eb6a-7d03-4f0f-96ab-d9f3773c0290Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM InferenceJiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie et al.NeurIPS 2025 · 74 citations
- TRACE: A Fast Transformer-based General-Purpose Lossless CompressorYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueWWW 2022 · 60 citations
- Towards Training Reproducible Deep Learning ModelsBoyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin et al.ICSE 2022 · 42 citations
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 31 citations
Related papers
- Lossless Compression with Probabilistic CircuitsAnji Liu, Stephan Mandt, Guy Van den BroeckICLR 2022 · 29 citations
- Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedKecheng Chen, Pingping Zhang, Hui Liu, Jie Liu et al.NeurIPS 2025 · 14 citations
- Language Models as Zero-shot Lossless Gradient Compressors: Towards General Neural Parameter Prior ModelsHui-Po Wang, Mario FritzNeurIPS 2024 · 5 citations
- Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood IntervalsChanghao PengCVPR 2025
- Progressive Neural Compression for Adaptive Image Offloading Under Timing ConstraintsRuiqi Wang, Hanyang Liu, Jiaming Qiu, Moran Xu et al.RTSS 2023 · 10 citations
