Efficient Mismatch-Tolerant Coding for Model-Driven Compression
Aviv Adler, Jennifer Tang
摘要
A central insight in lossless data compression is the close connection between probabilistic next- symbol prediction and efficient sequence compression, whereby predictive models can be combined with classical coding techniques to achieve strong compression performance. Applying this approach with powerful modern learned models, such as LLMs, has been shown to achieve markedly better compression than traditional techniques across a wide range of domains. However, significant practical challenges remain, including model non-determinism, in which a model produces different predictions on different machines despite identical parameters and inputs; such mismatches between the encoder and decoder can lead to complete decoding failure. Probability Matching Interval Coding (PMATIC) was recently introduced as a drop-in framework for mismatch-robust coding and shown to enable reliable compression and decompression in the presence of bounded prediction mismatch (Adler & Tang, 2026). In this work, we present a generalization of PMATIC that allows the incorpo- ration of tight theoretical results into the design and more flexible parameter optimization, resulting in substantial improvements in compression efficiency and robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt 等ICLR 2024 · 被引用 243 次
- Towards Training Reproducible Deep Learning ModelsBoyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin 等ICSE 2022 · 被引用 42 次
- Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedKecheng Chen, Pingping Zhang, Hui Liu, Jie Liu 等NeurIPS 2025 · 被引用 14 次
- Synchronizing Probabilities in Model-Driven Lossless CompressionAviv Adler, Jennifer TangICLR 2026 · 被引用 1 次
相关 Paper
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos 等NeurIPS 2024 · 被引用 29 次
- On the Out-of-distribution Generalization of Probabilistic Image ModellingMingtian Zhang, Andi Zhang, Steven McDonaghNeurIPS 2021 · 被引用 51 次
- Prompt-Guided Alignment with Information Bottleneck Makes Image Compression Also a RestorerXuelin Shen, Quan Liu, Jiayin Xu, Wenhan YangNeurIPS 2025
- CALLIC: Content Adaptive Learning for Lossless Image CompressionDaxin Li, Yuanchao Bai, Kai Wang, Junjun Jiang 等AAAI 2025 · 被引用 8 次
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
