Sequential Learning of Neural Networks for Prequential MDL
Jörg Bornschein, Yazhe Li, Marcus Hutter
Abstract
Minimum Description Length (MDL) provides a framework and an objective for principled model evaluation. It formalizes Occam's Razor and can be applied to data from non-stationary sources. In the prequential formulation of MDL, the objective is to minimize the cumulative next-step log-loss when sequentially going through the data and using previous observations for parameter estimation. It thus closely resembles a continual- or online-learning problem. In this study, we evaluate approaches for computing prequential description lengths for image classification datasets with neural networks. Considering the computational cost, we find that online-learning with rehearsal has favorable performance compared to the previously widely used block-wise estimation. We propose forward-calibration to better align the models predictions with the empirical observations and introduce replay-streams, a minibatch incremental training technique to efficiently implement approximate random replay while avoiding large in-memory replay buffers. As a result, we present description lengths for a suite of image classification datasets that improve upon previously reported results by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 528549f5-10b2-43dc-a343-e8d35a81d9ebCited by top-tier papers6
- Kalman Filter for Online Classification of Non-Stationary DataMichalis K. Titsias, Alexandre Galashov, Amal Rannen-Triki, Razvan Pascanu et al.ICLR 2024 · 14 citations
- Understanding Prompt Tuning and In-Context Learning via Meta-LearningTim Genewein, Kevin Li, Jordi Grau-Moya, Anian Ruoss et al.NeurIPS 2025 · 10 citations
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 7 citations
- Evaluating Representations with Readout Model SwitchingYazhe Li, Jörg Bornschein, Marcus HutterICLR 2023
- Towards a Formal Theory of Representational CompositionalityEric Elmoznino, Thomas Jiralerspong, Yoshua Bengio, Guillaume LajoieICML 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual DataZhipeng Cai, Ozan Sener, Vladlen KoltunICCV 2021 · 101 citations
Related papers
- A simple but strong baseline for online continual learning: Repeated Augmented RehearsalYaqian Zhang, Bernhard Pfahringer, Eibe Frank, Albert Bifet et al.NeurIPS 2022 · 13 citations
- Budgeted Online Continual Learning by Adaptive Layer Freezing and Frequency-based SamplingMinhyuk Seo, Hyunseo Koh, Jonghyun ChoiICLR 2025
- Online Class-Incremental Continual Learning with Adversarial Shapley ValueDongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner et al.AAAI 2021 · 262 citations
- Rethinking Momentum Knowledge Distillation in Online Continual LearningNicolas Michel, Maorong Wang, Ling Xiao, Toshihiko YamasakiICML 2024 · 26 citations
- F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental LearningHuiping Zhuang, Yuchen Liu, Run He, Kai Tong et al.NeurIPS 2024 · 17 citations
