Scalable Meta-Learning via Mixed-Mode Differentiation
Iurii Kemaev, Dan A. Calian, Luisa M. Zintgraf, Gregory Farquhar, Hado van Hasselt
Abstract
Gradient-based bilevel optimisation is a powerful technique with applications in hyperparameter optimisation, task adaptation, algorithm discovery, meta-learning more broadly, and beyond. It often requires differentiating through the gradientbased optimisation itself, leading to "gradientof-a-gradient" calculations with computationally expensive second-order and mixed derivatives. While modern automatic differentiation libraries provide a convenient way to write programs for calculating these derivatives, they oftentimes cannot fully exploit the specific structure of these problems out-of-the-box, leading to suboptimal performance. In this paper, we analyse such cases and propose Mixed-Flow Meta-Gradients, or MixFlow-MG -a practical algorithm that uses mixed-mode differentiation to construct more efficient and scalable computational graphs yielding over 10x memory and up to 25% wall-clock time improvements over standard implementations in modern meta-learning setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e09af9e5-a2ae-410a-86fe-e57ce2f0cebaCited by top-tier papers1
Ask how each one uses itBuilds on6
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Learning by Directional Gradient DescentDavid Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel et al.ICLR 2022 · 44 citations
- Making Scalable Meta Learning PracticalSang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger et al.NeurIPS 2023 · 28 citations
- Memory-Efficient Gradient Unrolling for Large-Scale Bi-level OptimizationQianli Shen, Yezhen Wang, Zhouhao Yang, Xiang Li et al.NeurIPS 2024 · 14 citations
Related papers
- Efficient Curvature-Aware Hypergradient Approximation for Bilevel OptimizationYouran Dong, Junfeng Yang, Wei Yao, Jin ZhangICML 2025
- Provably Faster Algorithms for Bilevel OptimizationJunjie Yang, Kaiyi Ji, Yingbin LiangNeurIPS 2021 · 175 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Betty: An Automatic Differentiation Library for Multilevel OptimizationSang Keun Choe, Willie Neiswanger, Pengtao Xie, Eric P. XingICLR 2023 · 5 citations
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 241 citations
