Chaos Meets Attention: Transformers for Large-Scale Dynamical Prediction
Yi He, Yiming Yang, Xiaoyuan Cheng, Hai Wang, Xiao Xue, Boli Chen, Yukun Hu
Abstract
Generating long-term trajectories of dissipative chaotic systems autoregressively is a highly challenging task. The inherent positive Lyapunov exponents amplify prediction errors over time. Many chaotic systems possess a crucial property -ergodicity on their attractors, which makes long-term prediction possible. State-of-the-art methods address ergodicity by preserving statistical properties using optimal transport techniques. However, these methods face scalability challenges due to the curse of dimensionality when matching distributions. To overcome this bottleneck, we propose a scalable transformerbased framework capable of stably generating long-term high-dimensional and high-resolution chaotic dynamics while preserving ergodicity. Our method is grounded in a physical perspective, revisiting the Von Neumann mean ergodic theorem to ensure the preservation of long-term statistics in the L 2 space. We introduce novel modifications to the attention mechanism, making the transformer architecture well-suited for learning large-scale chaotic systems. Compared to operator-based and transformer-based methods, our model achieves better performances across five metrics, from short-term prediction accuracy to long-term statistics. In addition to our methodological contributions, we introduce new chaotic system benchmarks: a machine learning dataset of 140k snapshots of turbulent channel flow and a processed high-dimensional Kolmogorov Flow dataset, along with various evaluation metrics for both short-and long-term performances. Both are well-suited for machine learning research on chaotic systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William GilpinICLR 2026 · 16 citations
- ChaosNexus: A Foundation Model for ODE-based Chaotic System Forecasting with Hierarchical Multi-scale AwarenessChang Liu, Bohao Zhao, Jingtao Ding, Yong LiICML 2026 · 1 citation
- MMPD-Bench: Bridging Multimodal Fission with Multi-Polarimetric Modalities DecompositionYi He, Zimo Zhao, Yiming Yang, Xiaoyuan Cheng et al.ICML 2026
- Tensor-Var: Efficient Four-Dimensional Variational Data AssimilationYiming Yang, Xiaoyuan Cheng, Daniel Giles, Sibo Cheng et al.ICML 2025
- Multi-Scale Wavelet Transformers for Operator Learning of Dynamical SystemsXuesong Wang, Michael Groom, Rafael Oliveira, He Zhao et al.ICML 2026
Builds on13
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
- GNOT: A General Neural Operator Transformer for Operator LearningZhongkai Hao, Zhengyi Wang, Hang Su, Chengyang Ying et al.ICML 2023 · 375 citations
- Multiwavelet-based Operator Learning for Differential EquationsGaurav Gupta, Xiongye Xiao, Paul BogdanNeurIPS 2021 · 355 citations
- Scalable Transformer for PDE Surrogate ModelingZijie Li, Dule Shu, Amir Barati FarimaniNeurIPS 2023 · 188 citations
Related papers
- Learning Chaotic Dynamics in Dissipative SystemsZongyi Li, Miguel Liu-Schiaffini, Nikola B. Kovachki, Kamyar Azizzadenesheli et al.NeurIPS 2022 · 62 citations
- Learning Chaos In A Linear WayXiaoyuan Cheng, Yi He, Yiming Yang, Xiao Xue et al.ICLR 2025
- DySLIM: Dynamics Stable Learning by Invariant Measure for Chaotic SystemsYair Schiff, Zhong Yi Wan, Jeffrey B. Parker, Stephan Hoyer et al.ICML 2024 · 30 citations
- Training neural operators to preserve invariant measures of chaotic attractorsRuoxi Jiang, Peter Y. Lu, Elena Orlova, Rebecca WillettNeurIPS 2023 · 59 citations
- When are dynamical systems learned from time series data statistically accurate?Jeongjin Park, Nicole Yang, Nisha ChandramoorthyNeurIPS 2024 · 17 citations
