Scaling Training Data with Lossy Image Compression
Katherine L. Mentzer, Andrea Montanari
摘要
Empirically-determined scaling laws have been broadly successful in predicting the evolution of large machine learning models with training data and number of parameters. As a consequence, they have been useful for optimizing the allocation of limited resources, most notably compute time. In certain applications, storage space is an important constraint, and data format needs to be chosen carefully as a consequence. Computer vision is a prominent example: images are inherently analog, but are always stored in a digital format using a finite number of bits. Given a dataset of digital images, the number of bits L to store each of them can be further reduced using lossy data compression. This, however, can degrade the quality of the model trained on such images, since each example has lower resolution. In order to capture this trade-off and optimize storage of training data, we propose a 'storage scaling law' that describes the joint evolution of test error with sample size and number of bits per image. We prove that this law holds within a stylized model for image compression, and verify it empirically on two computer vision tasks, extracting the relevant parameters. We then show that this law can be used to optimize the lossy compression level. At given storage, models trained on optimally compressed images present a significantly smaller test error with respect to models trained on the original data. Finally, we investigate the potential benefits of randomizing the compression level. * Granica 1 This of course holds under the condition that the model complexity is scaled simultaneously.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Scale Efficiently: Insights from Pretraining and Finetuning TransformersYi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus 等ICLR 2022 · 被引用 67 次
相关 Paper
- Unified Scaling Laws for Compressed RepresentationsAndrei Panferov, Alexandra Volkova, Ionut-Vlad Modoranu, Vage Egiazarian 等NeurIPS 2025 · 被引用 5 次
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon 等ICLR 2025
- AC-GC: Lossy Activation Compression with Guaranteed ConvergenceR. David Evans, Tor M. AamodtNeurIPS 2021 · 被引用 38 次
- A universal compression theory for lottery ticket hypothesis and neural scaling lawsHong-Yi Wang, Di Luo, Tomaso Poggio, Isaac L. Chuang 等ICLR 2026 · 被引用 2 次
- Optimal and Approximate Adaptive Stochastic QuantizationRan Ben-Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay VargaftikNeurIPS 2024 · 被引用 12 次
