Simple yet Effective: Low-Rank Spatial Attention for Neural Operators
Zherui Yang, Haiyang Xin, Tao Du, Ligang Liu
摘要
Neural operators have emerged as data-driven surrogates for solving partial differential equations (PDEs), and their success hinges on efficiently modeling the long-range, global coupling among spatial points induced by the underlying physics. In many PDE regimes, the induced global interaction kernels are empirically compressible, exhibiting rapid spectral decay that admits low-rank approximations. We leverage this observation to unify representative global mixing modules in neural operators under a shared low-rank template: compressing high-dimensional pointwise features into a compact latent space, processing global interactions within it, and reconstructing the global context back to spatial points. Guided by this view, we introduce Low-Rank Spatial Attention (LRSA) as a clean and direct instantiation of this template. Crucially, unlike prior approaches that often rely on non-standard aggregation or normalization modules, LRSA is built purely from standard Transformer primitives, i.e., attention, normalization, and feed-forward networks, yielding a concise block that is straightforward to implement and directly compatible with hardware-optimized kernels. In our experiments, such a simple construction is sufficient to achieve high accuracy, yielding an average error reduction of over 17% relative to second-best methods, while remaining stable and efficient in mixed-precision training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 被引用 1,175 次
相关 Paper
- Latent Neural Operator for Solving Forward and Inverse PDE ProblemsTian Wang, Chuang WangNeurIPS 2024 · 被引用 104 次
- Solving High-Dimensional PDEs with Latent Spectral ModelsHaixu Wu, Tengge Hu, Huakun Luo, Jianmin Wang 等ICML 2023 · 被引用 96 次
- AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural FieldsLouis Serrano, Thomas X. Wang, Etienne Le Naour, Jean-Noël Vittaut 等NeurIPS 2024 · 被引用 47 次
- Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEsSeungjun Lee, Taeil OhAAAI 2024 · 被引用 22 次
- Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics DiscoveryYue Yu, Ning Liu, Fei Lu, Tian Gao 等NeurIPS 2024 · 被引用 27 次
