Transolver Is a Linear Transformer: Revisiting Physics-Attention Through the Lens of Linear Attention
Wenjie Hu, Sidun Liu, Peng Qiao, Zhenglun Sun, Yong Dou
Abstract
Recent advances in Transformer-based Neural Operators have enabled significant progress in data-driven solvers for Partial Differential Equations (PDEs). Most current research has focused on reducing the quadratic complexity of attention to address the resulting low training and inference efficiency. Among these works, Transolver stands out as a representative method that introduces Physics-Attention to reduce computational costs. Physics-Attention projects grid points into slices for slice attention, then maps them back through deslicing. However, we observe that Physics-Attention can be reformulated as a special case of linear attention, and that the slice attention may even hurt the model performance. Based on these observations, we argue that its effectiveness primarily arises from the slice and deslice operations rather than interactions between slices. Building on this insight, we propose a two-step transformation to redesign Physics-Attention into a canonical linear attention, which we call Linear Attention Neural Operator (LinearNO). Our method achieves state-of-the-art performance on six standard PDE benchmarks, while reducing the number of parameters by an average of 40.0% and computational cost by 36.2%. Additionally, it delivers superior performance on two challenging, industrial-level datasets: AirfRANS and Shape-Net Car.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5ba433f-4dd9-4102-8b07-48ce166deeb8Cited by top-tier papers3
- Simple yet Effective: Low-Rank Spatial Attention for Neural OperatorsZherui Yang, Haiyang Xin, Tao Du, Ligang LiuICML 2026 · 2 citations
- Learning Laplacian Eigenspace with Mass-Aware Neural Operators on Point CloudsZherui Yang, Tao Du, Ligang LiuSIGGRAPH 2026
- Structure-Preserving Learning Improves Geometry Generalization in Neural PDEsBenjamin Shaffer, Shawn Koohy, Brooks Kinch, M. Ani Hsieh et al.ICML 2026
Builds on17
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 516 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
Related papers
- Transolver: A Fast Transformer Solver for PDEs on General GeometriesHaixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang et al.ICML 2024 · 228 citations
- Latent Neural Operator for Solving Forward and Inverse PDE ProblemsTian Wang, Chuang WangNeurIPS 2024 · 104 citations
- Neural Interpretable PDEs: Harmonizing Fourier Insights with Attention for Scalable and Interpretable Physics DiscoveryNing Liu, Yue YuICML 2025
- Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator LearningJunfeng Chen, Kailiang WuICML 2024 · 14 citations
- Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics DiscoveryYue Yu, Ning Liu, Fei Lu, Tian Gao et al.NeurIPS 2024 · 27 citations
