Improved Operator Learning by Orthogonal Attention
Zipeng Xiao, Zhongkai Hao, Bokai Lin, Zhijie Deng, Hang Su
Abstract
This work presents orthogonal attention for constructing neural operators to serve as surrogates to model the solutions of a family of Partial Differential Equations (PDEs). The motivation is that the kernel integral operator, which is usually at the core of neural operators, can be reformulated with orthonormal eigenfunctions. Inspired by the success of the neural approximation of eigenfunctions (Deng et al., 2022b), we opt to directly parameterize the involved eigenfunctions with flexible neural networks (NNs), based on which the input function is then transformed by the rule of kernel integral. Surprisingly, the resulting NN module bears a striking resemblance to regular attention mechanisms, albeit without softmax. Instead, it incorporates an orthogonalization operation that provides regularization during model training and helps mitigate overfitting, particularly in scenarios with limited data availability. In practice, the orthogonalization operation can be implemented with minimal additional overheads. Experiments on six standard neural operator benchmark datasets comprising both regular and irregular geometries show that our method can outperform competing baselines with decent margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ff7093d-b2e1-4977-8ce6-a16c84c21033Cited by top-tier papers30
- Transolver: A Fast Transformer Solver for PDEs on General GeometriesHaixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang et al.ICML 2024 · 228 citations
- Latent Neural Operator for Solving Forward and Inverse PDE ProblemsTian Wang, Chuang WangNeurIPS 2024 · 104 citations
- Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator LearningJunfeng Chen, Kailiang WuICML 2024 · 14 citations
- Spectral Convolutional Conditional Neural ProcessesPeiman Mohseni, Nick DuffieldNeurIPS 2025 · 10 citations
- Neural Green's FunctionsSeungwoo Yoo, Kyeongmin Yeo, Jisung Hwang, Minhyuk SungNeurIPS 2025 · 7 citations
Builds on9
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
Related papers
- Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics DiscoveryYue Yu, Ning Liu, Fei Lu, Tian Gao et al.NeurIPS 2024 · 27 citations
- SVD-NO: Learning PDE Solution Operators with SVD Integral KernelsNoam Koren, Ralf J. J. Mackenbach, Ruud J. G. van Sloun, Kira Radinsky et al.AAAI 2026
- Amortized Fourier Neural OperatorsZipeng Xiao, Siqi Kou, Zhongkai Hao, Bokai Lin et al.NeurIPS 2024 · 23 citations
- DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-TrainingZhongkai Hao, Chang Su, Songming Liu, Julius Berner et al.ICML 2024 · 107 citations
- NeuralEF: Deconstructing Kernels by Deep Neural NetworksZhijie Deng, Jiaxin Shi, Jun ZhuICML 2022 · 29 citations
