Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
Hong Wang, Haiyang Xin, Jie Wang, Xuanze Yang, Fei Zha, Huanshuo Dong, Yan Jiang
Abstract
Pre-training has proven effective in addressing data scarcity and performance limitations in solving PDE problems with neural operators. However, challenges remain due to the heterogeneity of PDE datasets in equation types, which leads to high errors in mixed training. Additionally, dense pre-training models that scale parameters by increasing network width or depth incur significant inference costs. To tackle these challenges, we propose a novel Mixture-of-Experts Pre-training Operator Transformer (MoE-POT), a sparse-activated architecture that scales parameters efficiently while controlling inference costs. Specifically, our model adopts a layer-wise router-gating network to dynamically select 4 routed experts from 16 expert networks during inference, enabling the model to focus on equationspecific features. Meanwhile, we also integrate 2 shared experts, aiming to capture common properties of PDE and reduce redundancy among routed experts. The final output is computed as the weighted average of the results from all activated experts. We pre-train models with parameters from 30M to 0.5B on 6 public PDE datasets. Our model with 90M activated parameters achieves up to a 40% reduction in zero-shot error compared with existing models with 120M activated parameters. Additionally, we conduct interpretability analysis, showing that dataset types can be inferred from router-gating network decisions, which validates the rationality and effectiveness of the MoE architecture 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-TrainingDengdi Sun, Xiaoya Zhou, Xiao Wang, Hao Si et al.CVPR 2026 · 2 citations
- Accelerating Eigenvalue Dataset Generation via Chebyshev Subspace FilterHong Wang, Jie Wang, Jian Luo, Huanshuo Dong et al.ICLR 2026 · 1 citation
- Learning-Guided Integration Contours Construction for Fast Large-Scale Generalized EigensolversYeqiu Chen, Ziyan Liu, Hong Wang, Lei LiuICML 2026 · 1 citation
- MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE TransformersZhao Yanshun, Xiaoyu Peng, Jiamin Jiang, Congcong Zhu et al.ICML 2026
- Origo: Interpretable Multi-physics PDE Foundation Model through Neural Operator SplittingLi Sun, Hongbo Lv, Zhikai Jiang, Zhongtian Sun et al.ICML 2026
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 516 citations
Related papers
- DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-TrainingZhongkai Hao, Chang Su, Songming Liu, Julius Berner et al.ICML 2024 · 107 citations
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-DesignRuisi Cai, Yeonju Ro, Geon-Woo Kim, Peihao Wang et al.NeurIPS 2024 · 21 citations
- Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable TransformersTianlong Chen, Zhenyu Zhang, Ajay Kumar Jaiswal, Shiwei Liu et al.ICLR 2023 · 6 citations
