Lune

ICLR2026Top-tier venue

A universal compression theory for lottery ticket hypothesis and neural scaling laws

Hong-Yi Wang, Di Luo, Tomaso Poggio, Isaac L. Chuang, Liu Ziyin

2026Year
2Citations

Abstract

When training large-scale models, the performance typically scales with the number of parameters and the dataset size according to a slow power law. A fundamental theoretical and practical question is whether comparable performance can be achieved with significantly smaller models and substantially less data. In this work, we provide a positive and constructive answer. We prove that a generic permutation-invariant function of dd objects can be asymptotically compressed into a function of polylog⁡d\operatorname{polylog} d objects with vanishing error, which is proved to be the optimal compression rate. This theorem yields two key implications: (Ia) a large neural network can be compressed to polylogarithmic width while preserving its learning dynamics; (Ib) a large dataset can be compressed to polylogarithmic size while leaving the loss landscape of the corresponding model unchanged. Implication (Ia) directly establishes a proof of the dynamical lottery ticket hypothesis, which states that any ordinary network can be strongly compressed such that the learning dynamics and result remain unchanged. (Ib) shows that a neural scaling law of the form L∼d−αL\sim d^{-\alpha} can be boosted to an arbitrarily fast power law decay, and ultimately to exp⁡(−α′dm)\exp(-\alpha' \sqrt[m]{d}).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 97e6ce9f-467e-4be9-95e2-6f8974d265f7

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines