CViT: Continuous Vision Transformer for Operator Learning
Sifan Wang, Jacob H. Seidman, Shyam Sankaran, Hanwen Wang, George J. Pappas, Paris Perdikaris
Abstract
Operator learning, which aims to approximate maps between infinite-dimensional function spaces, is an important area in scientific machine learning with applications across various physical domains. Here we introduce the Continuous Vision Transformer (CViT), a novel neural operator architecture that leverages advances in computer vision to address challenges in learning complex physical systems. CViT combines a vision transformer encoder, a novel grid-based coordinate embedding, and a query-wise cross-attention mechanism to effectively capture multi-scale dependencies. This design allows for flexible output representations and consistent evaluation at arbitrary resolutions. We demonstrate CViT's effectiveness across a diverse range of partial differential equation (PDE) systems, including fluid dynamics, climate modeling, and reaction-diffusion processes. Our comprehensive experiments show that CViT achieves state-of-the-art performance on multiple benchmarks, often surpassing larger foundation models, even without extensive pretraining and roll-out fine-tuning. Taken together, CViT exhibits robust handling of discontinuous solutions, multi-scale features, and intricate spatio-temporal dynamics. Our contributions can be viewed as a significant step towards adapting advanced computer vision architectures for building more flexible and accurate machine learning models in the physical sciences. All data and code are publicly available at https://github.com/PredictiveIntelligenceLab/cvit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- ENMA: Tokenwise Autoregression for Continuous Neural PDE OperatorsArmand Kassaï Koupaï, Lise Le Boudec, Louis Serrano, Patrick GallinariNeurIPS 2025 · 9 citations
- MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale AttentionPedro M. P. Curvo, Jan-Willem van de Meent, Maksim ZhdanovCVPR 2026 · 3 citations
- Overtone: Cyclic Patch Modulation for Clean, Efficient, and Flexible Physics EmulatorsPayel Mukhopadhyay, Michael McCabe, Ruben Ohana, Miles D. CranmerICLR 2026 · 3 citations
- PhysicsCorrect: A Training-Free Approach for Stable Neural PDE SimulationsXinquan Huang, Paris PerdikarisAAAI 2026 · 3 citations
- Adaptive Physics Transformer with Fused Global-Local Attention for Subsurface Energy SystemsXin Ju, Hadrian Fung, Yuyan Zhang, Carl Jacquemyn et al.ICML 2026 · 2 citations
Builds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
Related papers
- Geometry Aware Operator Transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domainsShizheng Wen, Arsh Kumbhat, Levi E. Lingsch, Sepehr Mousavi et al.NeurIPS 2025 · 73 citations
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
- NOMAD: Nonlinear Manifold Decoders for Operator LearningJacob H. Seidman, Georgios Kissas, Paris Perdikaris, George J. PappasNeurIPS 2022 · 125 citations
- RIGNO: A Graph-based Framework For Robust And Accurate Operator Learning For PDEs On Arbitrary DomainsSepehr Mousavi, Shizheng Wen, Levi E. Lingsch, Maximilian Herde et al.NeurIPS 2025 · 31 citations
- Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEsMd. Ashiqur Rahman, Robert Joseph George, Mogab Elleithy, Daniel V. Leibovici et al.NeurIPS 2024 · 79 citations
