Choose a Transformer: Fourier or Galerkin
Shuhao Cao
Abstract
In this paper, we apply the self-attention from the state-of-the-art Transformer in Attention Is All You Need [88] for the first time to a data-driven operator learning problem related to partial differential equations. An effort is put together to explain the heuristics of, and to improve the efficacy of the attention mechanism. By employing the operator approximation theory in Hilbert spaces, it is demonstrated for the first time that the softmax normalization in the scaled dot-product attention is sufficient but not necessary. Without softmax, the approximation capacity of a linearized Transformer variant can be proved to be comparable to a Petrov-Galerkin projection layer-wise, and the estimate is independent with respect to the sequence length. A new layer normalization scheme mimicking the Petrov-Galerkin projection is proposed to allow a scaling to propagate through attention layers, which helps the model achieve remarkable accuracy in operator learning tasks with unnormalized data. Finally, we present three operator learning experiments, including the viscid Burgers' equation, an interface Darcy flow, and an inverse interface coefficient identification problem. The newly proposed simple attention-based operator learner, Galerkin Transformer, shows significant improvements in both training cost and evaluation accuracy over its softmax-normalized counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5232f272-0444-496b-a526-a4eee2f838abCited by top-tier papers91
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
- Convolutional Neural Operators for robust and accurate learning of PDEsBogdan Raonic, Roberto Molinaro, Tim De Ryck, Tobias Rohner et al.NeurIPS 2023 · 292 citations
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli et al.NeurIPS 2024 · 235 citations
- Transolver: A Fast Transformer Solver for PDEs on General GeometriesHaixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang et al.ICML 2024 · 228 citations
- Scalable Transformer for PDE Surrogate ModelingZijie Li, Dule Shu, Amir Barati FarimaniNeurIPS 2023 · 188 citations
Builds on18
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
Related papers
- Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator LearningJunfeng Chen, Kailiang WuICML 2024 · 14 citations
- Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionWenjie Hu, Sidun Liu, Peng Qiao, Zhenglun Sun et al.AAAI 2026 · 2 citations
- Geometry Aware Operator Transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domainsShizheng Wen, Arsh Kumbhat, Levi E. Lingsch, Sepehr Mousavi et al.NeurIPS 2025 · 73 citations
- Exploring Efficient Partial Differential Equation Solution Using Speed Galerkin TransformerXun Wang, Zeyang Zhu, Xiangyu Meng, Tao SongSC 2024
- Multiwavelet-based Operator Learning for Differential EquationsGaurav Gupta, Xiongye Xiao, Paul BogdanNeurIPS 2021 · 355 citations
