DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training
Zhongkai Hao, Chang Su, Songming Liu, Julius Berner, Chengyang Ying, Hang Su, Anima Anandkumar, Jian Song, Jun Zhu
Abstract
Pre-training has been investigated to improve the efficiency and performance of training neural operators in data-scarce settings. However, it is largely in its infancy due to the inherent complexity and diversity, such as long trajectories, multiple scales and varying dimensions of partial differential equations (PDEs) data. In this paper, we present a new auto-regressive denoising pre-training strategy, which allows for more stable and efficient pre-training on PDE data and generalizes to various downstream tasks. Moreover, by designing a flexible and scalable model architecture based on Fourier attention, we can easily scale up the model for large-scale pre-training. We train our PDE foundation model with up to 0.5B parameters on 10+ PDE datasets with more than 100k trajectories. Extensive experiments show that we achieve SOTA on these benchmarks and validate the strong generalizability of our model to significantly enhance performance on diverse downstream PDE tasks like 3D data. Code is available at https://github.com/thu-ml/DPOT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2c296f3-2c72-458a-aa28-6218227f7799Cited by top-tier papers51
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli et al.NeurIPS 2024 · 235 citations
- Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEsMd. Ashiqur Rahman, Robert Joseph George, Mogab Elleithy, Daniel V. Leibovici et al.NeurIPS 2024 · 79 citations
- Data-Efficient Operator Learning via Unsupervised Pretraining and In-Context LearningWuyang Chen, Jialin Song, Pu Ren, Shashank Subramanian et al.NeurIPS 2024 · 41 citations
- RIGNO: A Graph-based Framework For Robust And Accurate Operator Learning For PDEs On Arbitrary DomainsSepehr Mousavi, Shizheng Wen, Levi E. Lingsch, Maximilian Herde et al.NeurIPS 2025 · 31 citations
- Universal Physics Transformers: A Framework For Efficiently Scaling Neural OperatorsBenedikt Alkin, Andreas Fürst, Simon Schmid, Lukas Gruber et al.NeurIPS 2024 · 23 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Geometry-Informed Neural Operator for Large-Scale 3D PDEsZongyi Li, Nikola B. Kovachki, Christopher B. Choy, Boyi Li et al.NeurIPS 2023 · 461 citations
Related papers
- NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-TrainingDengdi Sun, Xiaoya Zhou, Xiao Wang, Hao Si et al.CVPR 2026 · 2 citations
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-TrainingHong Wang, Haiyang Xin, Jie Wang, Xuanze Yang et al.NeurIPS 2025 · 15 citations
- Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEsSeungjun Lee, Taeil OhAAAI 2024 · 22 citations
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
- Spectral-Inspired Neural Operator Learning with Limited Data and Unknown PhysicsHan Wan, Rui Zhang, Hao SunKDD 2026 · 1 citation
