MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE Transformers
Zhao Yanshun, Xiaoyu Peng, Jiamin Jiang, Congcong Zhu, Jingrun Chen
Abstract
Conventional patchified Transformers operate on uniform spatial partitions, distributing computational effort evenly across the domain irrespective of local features. This inflexible tokenization scheme is inherently limited in its ability to efficiently represent and process solutions to complex PDEs. To address this, we propose MeshTok, an adaptive mesh refinement (AMR)-inspired tokenization and sequence modeling framework. This method selectively refines spatial regions exhibiting sharp gradients, transient features, or multiscale structures, generating a heterogeneous set of multiscale tokens defined on a fixed simulation grid. These tokens are processed within a unified Transformer sequence, enabling the model to simultaneously capture coarse-grained global context and fine-grained local details without requiring specialized architectural components. Although adaptive refinement moderately increases token count, it promotes a more targeted allocation of computational resources to physically informative regions, which we view as a practical inductive bias rather than a formal optimality guarantee. Experimental evaluations across multiple PDE families and benchmark datasets demonstrate that MeshTok consistently improves the efficiency-accuracy trade-off compared to uniform-grid baselines. This suggests adaptive multiscale tokenization as a scalable and generalizable design principle for neural PDE modeling. Code is available at https://github.com/ SCAILab-USTC/MeshTok .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61169ad6-c680-451d-a873-42d9f51017c7Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
Related papers
- MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale AttentionPedro M. P. Curvo, Jan-Willem van de Meent, Maksim ZhdanovCVPR 2026 · 3 citations
- AMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid SimulationZeyi Xu, Jinfan Liu, Kuangxu Chen, Ye Chen et al.CVPR 2025
- Adaptive Physics Transformer with Fused Global-Local Attention for Subsurface Energy SystemsXin Ju, Hadrian Fung, Yuyan Zhang, Carl Jacquemyn et al.ICML 2026 · 2 citations
- SpiderSolver: A Geometry-Aware Transformer for Solving PDEs on Complex GeometriesKai Qi, Fan Wang, Zhewen Dong, Jian SunNeurIPS 2025 · 3 citations
- Adaptive Patching for High-resolution Image Segmentation with TransformersEnzhi Zhang, Isaac Lyngaas, Peng Chen, Xiao Wang et al.SC 2024 · 6 citations
