Scaling Training Data with Lossy Image Compression
Katherine L. Mentzer, Andrea Montanari
Abstract
Empirically-determined scaling laws have been broadly successful in predicting the evolution of large machine learning models with training data and number of parameters. As a consequence, they have been useful for optimizing the allocation of limited resources, most notably compute time. In certain applications, storage space is an important constraint, and data format needs to be chosen carefully as a consequence. Computer vision is a prominent example: images are inherently analog, but are always stored in a digital format using a finite number of bits. Given a dataset of digital images, the number of bits L to store each of them can be further reduced using lossy data compression. This, however, can degrade the quality of the model trained on such images, since each example has lower resolution. In order to capture this trade-off and optimize storage of training data, we propose a 'storage scaling law' that describes the joint evolution of test error with sample size and number of bits per image. We prove that this law holds within a stylized model for image compression, and verify it empirically on two computer vision tasks, extracting the relevant parameters. We then show that this law can be used to optimize the lossy compression level. At given storage, models trained on optimally compressed images present a significantly smaller test error with respect to models trained on the original data. Finally, we investigate the potential benefits of randomizing the compression level. * Granica 1 This of course holds under the condition that the model complexity is scaled simultaneously.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 171 citations
- Scale Efficiently: Insights from Pretraining and Finetuning TransformersYi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus et al.ICLR 2022 · 67 citations
Related papers
- Unified Scaling Laws for Compressed RepresentationsAndrei Panferov, Alexandra Volkova, Ionut-Vlad Modoranu, Vage Egiazarian et al.NeurIPS 2025 · 5 citations
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon et al.ICLR 2025
- AC-GC: Lossy Activation Compression with Guaranteed ConvergenceR. David Evans, Tor M. AamodtNeurIPS 2021 · 38 citations
- A universal compression theory for lottery ticket hypothesis and neural scaling lawsHong-Yi Wang, Di Luo, Tomaso Poggio, Isaac L. Chuang et al.ICLR 2026 · 2 citations
- Optimal and Approximate Adaptive Stochastic QuantizationRan Ben-Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay VargaftikNeurIPS 2024 · 12 citations
