Simple yet Effective: Low-Rank Spatial Attention for Neural Operators
Zherui Yang, Haiyang Xin, Tao Du, Ligang Liu
Abstract
Neural operators have emerged as data-driven surrogates for solving partial differential equations (PDEs), and their success hinges on efficiently modeling the long-range, global coupling among spatial points induced by the underlying physics. In many PDE regimes, the induced global interaction kernels are empirically compressible, exhibiting rapid spectral decay that admits low-rank approximations. We leverage this observation to unify representative global mixing modules in neural operators under a shared low-rank template: compressing high-dimensional pointwise features into a compact latent space, processing global interactions within it, and reconstructing the global context back to spatial points. Guided by this view, we introduce Low-Rank Spatial Attention (LRSA) as a clean and direct instantiation of this template. Crucially, unlike prior approaches that often rely on non-standard aggregation or normalization modules, LRSA is built purely from standard Transformer primitives, i.e., attention, normalization, and feed-forward networks, yielding a concise block that is straightforward to implement and directly compatible with hardware-optimized kernels. In our experiments, such a simple construction is sufficient to achieve high accuracy, yielding an average error reduction of over 17% relative to second-best methods, while remaining stable and efficient in mixed-precision training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 454e85c0-b4bf-4cee-aeba-0dcb571f3e98Cited by top-tier papers1
Ask how each one uses itBuilds on23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
Related papers
- Latent Neural Operator for Solving Forward and Inverse PDE ProblemsTian Wang, Chuang WangNeurIPS 2024 · 104 citations
- Solving High-Dimensional PDEs with Latent Spectral ModelsHaixu Wu, Tengge Hu, Huakun Luo, Jianmin Wang et al.ICML 2023 · 96 citations
- AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural FieldsLouis Serrano, Thomas X. Wang, Etienne Le Naour, Jean-Noël Vittaut et al.NeurIPS 2024 · 47 citations
- Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEsSeungjun Lee, Taeil OhAAAI 2024 · 22 citations
- Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics DiscoveryYue Yu, Ning Liu, Fei Lu, Tian Gao et al.NeurIPS 2024 · 27 citations
