MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
Pedro M. P. Curvo, Jan-Willem van de Meent, Maksim Zhdanov
Abstract
A key scalability challenge in neural solvers for industrialscale physics simulations is efficiently capturing both finegrained local interactions and long-range global dependencies across millions of spatial elements. We introduce the Multi-Scale Patch Transformer (MSPT) 1 , an architecture that combines local point attention within patches with global attention to coarse patch-level representations. To partition the input domain into spatially-coherent patches, we employ ball trees, which handle irregular geometries efficiently. This dual-scale design enables MSPT to scale to millions of points on a single GPU. We validate our method on standard PDE benchmarks (elasticity, plasticity, fluid dynamics, porous flow) and large-scale aerodynamic datasets (ShapeNet-Car, Ahmed-ML), achieving state-of-the-art accuracy with substantially lower memory footprint and computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26078326-9e42-4255-a916-013afc87f1cbBuilds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- Learning Mesh-Based Simulation with Graph NetworksTobias Pfaff, Meire Fortunato, Alvaro Sanchez-Gonzalez, Peter W. BattagliaICLR 2021 · 1,175 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 516 citations
Related papers
- Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical SystemsMaksim Zhdanov, Max Welling, Jan-Willem van de MeentICML 2025
- Transolver++: An Accurate Neural Solver for PDEs on Million-Scale GeometriesHuakun Luo, Haixu Wu, Hang Zhou, Lanxiang Xing et al.ICML 2025
- Transolver: A Fast Transformer Solver for PDEs on General GeometriesHaixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang et al.ICML 2024 · 228 citations
- A Hierarchical Spatial Transformer for Massive Point Samples in Continuous SpaceWenchong He, Zhe Jiang, Tingsong Xiao, Zelin Xu et al.NeurIPS 2023 · 20 citations
- Transolver-3: Scaling Up Transformer Solvers to Industrial-Scale GeometriesHang Zhou, Haixu Wu, Haonan Shangguan, Yuezhou Ma et al.ICML 2026
