Choose a Transformer: Fourier or Galerkin
Shuhao Cao
摘要
In this paper, we apply the self-attention from the state-of-the-art Transformer in Attention Is All You Need [88] for the first time to a data-driven operator learning problem related to partial differential equations. An effort is put together to explain the heuristics of, and to improve the efficacy of the attention mechanism. By employing the operator approximation theory in Hilbert spaces, it is demonstrated for the first time that the softmax normalization in the scaled dot-product attention is sufficient but not necessary. Without softmax, the approximation capacity of a linearized Transformer variant can be proved to be comparable to a Petrov-Galerkin projection layer-wise, and the estimate is independent with respect to the sequence length. A new layer normalization scheme mimicking the Petrov-Galerkin projection is proposed to allow a scaling to propagate through attention layers, which helps the model achieve remarkable accuracy in operator learning tasks with unnormalized data. Finally, we present three operator learning experiments, including the viscid Burgers' equation, an interface Darcy flow, and an inverse interface coefficient identification problem. The newly proposed simple attention-based operator learner, Galerkin Transformer, shows significant improvements in both training cost and evaluation accuracy over its softmax-normalized counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper91
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying 等ICML 2023 · 被引用 375 次
- Convolutional Neural Operators for robust and accurate learning of PDEsBogdan Raonic, Roberto Molinaro, Tim De Ryck, Tobias Rohner 等NeurIPS 2023 · 被引用 292 次
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli 等NeurIPS 2024 · 被引用 235 次
- Transolver: A Fast Transformer Solver for PDEs on General GeometriesHaixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang 等ICML 2024 · 被引用 228 次
- Scalable Transformer for PDE Surrogate ModelingZijie Li, Dule Shu, Amir Barati FarimaniNeurIPS 2023 · 被引用 188 次
它引用的顶会 Paper18
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
相关 Paper
- Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator LearningJunfeng Chen, Kailiang WuICML 2024 · 被引用 14 次
- Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear AttentionWenjie Hu, Sidun Liu, Peng Qiao, Zhenglun Sun 等AAAI 2026 · 被引用 2 次
- Geometry Aware Operator Transformer as an efficient and accurate neural surrogate for PDEs on arbitrary domainsShizheng Wen, Arsh Kumbhat, Levi E. Lingsch, Sepehr Mousavi 等NeurIPS 2025 · 被引用 73 次
- Exploring Efficient Partial Differential Equation Solution Using Speed Galerkin TransformerXun Wang, Zeyang Zhu, Xiangyu Meng, Tao SongSC 2024
- Multiwavelet-based Operator Learning for Differential EquationsGaurav Gupta, Xiongye Xiao, Paul BogdanNeurIPS 2021 · 被引用 355 次
