Making Scalable Meta Learning Practical
Sang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger, Pengtao Xie, Emma Strubell, Eric P. Xing
Abstract
Despite its flexibility to learn diverse inductive biases in machine learning programs, meta learning (i.e., learning to learn) has long been recognized to suffer from poor scalability due to its tremendous compute/memory costs, training instability, and a lack of efficient distributed training support. In this work, we focus on making scalable meta learning practical by introducing SAMA, which combines advances in both implicit differentiation algorithms and systems. Specifically, SAMA is designed to flexibly support a broad range of adaptive optimizers in the base level of meta learning programs, while reducing computational burden by avoiding explicit computation of second-order gradient information, and exploiting efficient distributed training techniques implemented for first-order gradients. Evaluated on multiple large-scale meta learning benchmarks, SAMA showcases up to 1.7/4.8x increase in throughput and 2.0/3.8x decrease in memory consumption respectively on single-/multi-GPU setups compared to other baseline meta learning algorithms. Furthermore, we show that SAMA-based data optimization leads to consistent improvements in text classification accuracy with BERT and RoBERTa large language models, and achieves state-of-the-art results in both small- and large-scale data pruning on image classification tasks, demonstrating the practical applicability of scalable meta learning across language and vision domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai et al.ICML 2026 · 24 citations
- Memory-Efficient Gradient Unrolling for Large-Scale Bi-level OptimizationQianli Shen, Yezhen Wang, Zhouhao Yang, Xiang Li et al.NeurIPS 2024 · 14 citations
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient TracingZhe Li, Wei Zhao, Yige Li, Jun SunICLR 2026 · 4 citations
- Efficient Bilevel Source Mask OptimizationGuojin Chen, Hongquan He, Peng Xu, Hao Geng et al.DAC 2024 · 4 citations
- TANDEM: Bi-Level Data Mixture Optimization with Twin NetworksJiaxing Wang, Deping Xiang, Jin Xu, Mingyang Yi et al.NeurIPS 2025 · 3 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- Memory Efficient Meta-Learning with Large ImagesJohn Bronskill, Daniela Massiceti, Massimiliano Patacchiola, Katja Hofmann et al.NeurIPS 2021 · 21 citations
- Towards Efficient Low-Order Hybrid Optimizer for Language Model Fine-TuningMinping Chen, You-Liang Huang, Zeyi WenAAAI 2025 · 6 citations
- Memory-Reduced Meta-Learning with Guaranteed ConvergenceHonglin Yang, Ji Ma, Xiao YuAAAI 2025 · 1 citation
- HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order OptimizationHuaqin Zhao, Jiaxi Li, Yi Pan, Shizhe Liang et al.EMNLP 2025
- Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-TrainingReza Shirkavand, Peiran Yu, Qi He, Heng HuangNeurIPS 2025 · 6 citations
