Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEs
Md. Ashiqur Rahman, Robert Joseph George, Mogab Elleithy, Daniel V. Leibovici, Zongyi Li, Boris Bonev, Colin White, Julius Berner, Raymond A. Yeh, Jean Kossaifi, Kamyar Azizzadenesheli, Animashree Anandkumar
Abstract
Existing neural operator architectures face challenges when solving multiphysics problems with coupled partial differential equations (PDEs) due to complex geometries, interactions between physical variables, and the limited amounts of high-resolution training data. To address these issues, we propose Codomain Attention Neural Operator (CoDA-NO), which tokenizes functions along the codomain or channel space, enabling self-supervised learning or pretraining of multiple PDE systems. Specifically, we extend positional encoding, self-attention, and normalization layers to function spaces. CoDA-NO can learn representations of different PDE systems with a single model. We evaluate CoDA-NO's potential as a backbone for learning multiphysics PDEs over multiple systems by considering few-shot learning settings. On complex downstream tasks with limited data, such as fluid flow simulations, fluid-structure interactions, and Rayleigh-Bénard convection, we found CoDA-NO to outperform existing methods by over 36%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Neural Operators with Localized Integral and Differential KernelsMiguel Liu-Schiaffini, Julius Berner, Boris Bonev, Thorsten Kurth et al.ICML 2024 · 63 citations
- Walrus: A Cross-domain Foundation Model for Continuum DynamicsMichael McCabe, Payel Mukhopadhyay, Tanya Marwah, Bruno Régaldo-Saint Blancard et al.ICML 2026 · 23 citations
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William GilpinICLR 2026 · 16 citations
- EddyFormer: Accelerated Neural Simulations of Three-Dimensional Turbulence at ScaleYiheng Du, Aditi S. KrishnapriyanNeurIPS 2025 · 8 citations
- Generative Adaptation of Dynamics to Environmental Shifts via Weight-space DiffusionRuikun Li, Huandong Wang, Jingtao Ding, Yuan Yuan et al.ICML 2026 · 4 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics DiscoveryYue Yu, Ning Liu, Fei Lu, Tian Gao et al.NeurIPS 2024 · 27 citations
- Latent Neural Operator for Solving Forward and Inverse PDE ProblemsTian Wang, Chuang WangNeurIPS 2024 · 104 citations
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli et al.NeurIPS 2024 · 235 citations
- Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics DiscoveryNing Liu, Yue YuICML 2025
- CViT: Continuous Vision Transformer for Operator LearningSifan Wang, Jacob H. Seidman, Shyam Sankaran, Hanwen Wang et al.ICLR 2025
