Efficient Mismatch-Tolerant Coding for Model-Driven Compression
Aviv Adler, Jennifer Tang
Abstract
A central insight in lossless data compression is the close connection between probabilistic next- symbol prediction and efficient sequence compression, whereby predictive models can be combined with classical coding techniques to achieve strong compression performance. Applying this approach with powerful modern learned models, such as LLMs, has been shown to achieve markedly better compression than traditional techniques across a wide range of domains. However, significant practical challenges remain, including model non-determinism, in which a model produces different predictions on different machines despite identical parameters and inputs; such mismatches between the encoder and decoder can lead to complete decoding failure. Probability Matching Interval Coding (PMATIC) was recently introduced as a drop-in framework for mismatch-robust coding and shown to enable reliable compression and decompression in the presence of bounded prediction mismatch (Adler & Tang, 2026). In this work, we present a generalization of PMATIC that allows the incorpo- ration of tight theoretical results into the design and more flexible parameter optimization, resulting in substantial improvements in compression efficiency and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
- Towards Training Reproducible Deep Learning ModelsBoyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin et al.ICSE 2022 · 42 citations
- Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedKecheng Chen, Pingping Zhang, Hui Liu, Jie Liu et al.NeurIPS 2025 · 14 citations
- Synchronizing Probabilities in Model-Driven Lossless CompressionAviv Adler, Jennifer TangICLR 2026 · 1 citation
Related papers
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos et al.NeurIPS 2024 · 29 citations
- On the Out-of-distribution Generalization of Probabilistic Image ModellingMingtian Zhang, Andi Zhang, Steven McDonaghNeurIPS 2021 · 51 citations
- Prompt-Guided Alignment with Information Bottleneck Makes Image Compression Also a RestorerXuelin Shen, Quan Liu, Jiayin Xu, Wenhan YangNeurIPS 2025
- CALLIC: Content Adaptive Learning for Lossless Image CompressionDaxin Li, Yuanchao Bai, Kai Wang, Junjun Jiang et al.AAAI 2025 · 8 citations
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen et al.ICLR 2026 · 6 citations
