Scaling Down Deep Learning with MNIST-1D
Samuel Greydanus, Dmitry Kobak
Abstract
Although deep learning models have taken on commercial and political relevance, key aspects of their training and operation remain poorly understood. This has sparked interest in science of deep learning projects, many of which require large amounts of time, money, and electricity. But how much of this research really needs to occur at scale? In this paper, we introduce MNIST-1D: a minimalist, procedurally generated, low-memory, and low-compute alternative to classic deep learning benchmarks. Although the dimensionality of MNIST-1D is only 40 and its default training set size only 4000, MNIST-1D can be used to study inductive biases of different deep architectures, find lottery tickets, observe deep double descent, metalearn an activation function, and demonstrate guillotine regularization in self-supervised learning. All these experiments can be conducted on a GPU or often even on a CPU within minutes, allowing for fast prototyping, educational use cases, and cutting-edge research on a low budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe44d8b2-2fc8-409f-8eb9-229e27d94f52Cited by top-tier papers3
- Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondAlan Jeffares, Alicia Curth, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Efficient Source-Free Time-Series Adaptation via Parameter Subspace DisentanglementGaurav Patel, Christopher Michael Sandino, Behrooz Mahasseni, Ellen L. Zippi et al.ICLR 2025
- Deep Learning with Learnable Product-Structured ActivationsSaanjali Maharaj, Prasanth B. NairICLR 2026
Builds on2
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
Related papers
- Convolutional and Residual Networks Provably Contain Lottery TicketsRebekka BurkholzICML 2022 · 18 citations
- How many degrees of freedom do we need to train deep networks: a loss landscape perspectiveBrett W. Larsen, Stanislav Fort, Nic Becker, Surya GanguliICLR 2022 · 33 citations
- Sanity Checks for Lottery Tickets: Does Your Winning Ticket Really Win the Jackpot?Xiaolong Ma, Geng Yuan, Xuan Shen, Tianlong Chen et al.NeurIPS 2021 · 73 citations
- Small Data, Big Decisions: Model Selection in the Small-Data RegimeJörg Bornschein, Francesco Visin, Simon OsinderoICML 2020 · 48 citations
- Scaling MLPs: A Tale of Inductive BiasGregor Bachmann, Sotiris Anagnostidis, Thomas HofmannNeurIPS 2023 · 71 citations
