Sequential Learning of Neural Networks for Prequential MDL
Jörg Bornschein, Yazhe Li, Marcus Hutter
摘要
Minimum Description Length (MDL) provides a framework and an objective for principled model evaluation. It formalizes Occam's Razor and can be applied to data from non-stationary sources. In the prequential formulation of MDL, the objective is to minimize the cumulative next-step log-loss when sequentially going through the data and using previous observations for parameter estimation. It thus closely resembles a continual- or online-learning problem. In this study, we evaluate approaches for computing prequential description lengths for image classification datasets with neural networks. Considering the computational cost, we find that online-learning with rehearsal has favorable performance compared to the previously widely used block-wise estimation. We propose forward-calibration to better align the models predictions with the empirical observations and introduce replay-streams, a minibatch incremental training technique to efficiently implement approximate random replay while avoiding large in-memory replay buffers. As a result, we present description lengths for a suite of image classification datasets that improve upon previously reported results by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Kalman Filter for Online Classification of Non-Stationary DataMichalis K. Titsias, Alexandre Galashov, Amal Rannen-Triki, Razvan Pascanu 等ICLR 2024 · 被引用 14 次
- Understanding Prompt Tuning and In-Context Learning via Meta-LearningTim Genewein, Kevin Li, Jordi Grau-Moya, Anian Ruoss 等NeurIPS 2025 · 被引用 10 次
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 被引用 7 次
- Evaluating Representations with Readout Model SwitchingYazhe Li, Jörg Bornschein, Marcus HutterICLR 2023
- Towards a Formal Theory of Representational CompositionalityEric Elmoznino, Thomas Jiralerspong, Yoshua Bengio, Guillaume LajoieICML 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 被引用 288 次
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual DataZhipeng Cai, Ozan Sener, Vladlen KoltunICCV 2021 · 被引用 101 次
相关 Paper
- A simple but strong baseline for online continual learning: Repeated Augmented RehearsalYaqian Zhang, Bernhard Pfahringer, Eibe Frank, Albert Bifet 等NeurIPS 2022 · 被引用 13 次
- Budgeted Online Continual Learning by Adaptive Layer Freezing and Frequency-based SamplingMinhyuk Seo, Hyunseo Koh, Jonghyun ChoiICLR 2025
- Online Class-Incremental Continual Learning with Adversarial Shapley ValueDongsub Shim, Zheda Mai, Jihwan Jeong, Scott Sanner 等AAAI 2021 · 被引用 262 次
- Rethinking Momentum Knowledge Distillation in Online Continual LearningNicolas Michel, Maorong Wang, Ling Xiao, Toshihiko YamasakiICML 2024 · 被引用 26 次
- F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental LearningHuiping Zhuang, Yuchen Liu, Run He, Kai Tong 等NeurIPS 2024 · 被引用 17 次
