Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
Shanda Li, Shinjae Yoo, Yiming Yang
Abstract
Fourier Neural Operators (FNOs) offer a principled approach for solving complex partial differential equations (PDEs). However, scaling them to handle more complex PDEs requires increasing the number of Fourier modes, which significantly expands the number of model parameters and makes hyperparameter tuning computationally impractical. To address this, we introduce µTransfer-FNO, a zero-shot hyperparameter transfer technique that enables optimal configurations, tuned on smaller FNOs, to be directly applied to billion-parameter FNOs without additional tuning. Building on the Maximal Update Parametrization (µP) framework, we mathematically derive a parametrization scheme that facilitates the transfer of optimal hyperparameters across models with different numbers of Fourier modes in FNOs, which is validated through extensive experiments on various PDEs. Our empirical study shows that µTransfer-FNO reduces computational cost for tuning hyperparameters on large FNOs while maintaining or improving accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c36789e3-ae7d-409b-bfc6-66bb081b4fc0Builds on16
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Multipole Graph Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.NeurIPS 2020 · 569 citations
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 516 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Generic bounds on the approximation error for physics-informed (and) operator learningTim De Ryck, Siddhartha MishraNeurIPS 2022 · 93 citations
Related papers
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Amortized Fourier Neural OperatorsZipeng Xiao, Siqi Kou, Zhongkai Hao, Bokai Lin et al.NeurIPS 2024 · 23 citations
- Extending Fourier Neural Operators for Modeling Parameterized and Coupled PDEsCheng Jing, Uvini Balasuriya Mudiyanselage, Abhishek Verma, Kallol Bera et al.ICLR 2026 · 1 citation
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Understanding the Mechanisms of Fast Hyperparameter TransferNikhil Ghosh, Denny Wu, Alberto BiettiICLR 2026 · 8 citations
