Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
Shashank Subramanian, Peter Harrington, Kurt Keutzer, Wahid Bhimji, Dmitriy Morozov, Michael W. Mahoney, Amir Gholami
摘要
Pre-trained machine learning (ML) models have shown great performance for a wide range of applications, in particular in natural language processing (NLP) and computer vision (CV). Here, we study how pre-training could be used for scientific machine learning (SciML) applications, specifically in the context of transfer learning. We study the transfer behavior of these models as (i) the pre-trained model size is scaled, (ii) the downstream training dataset size is scaled, (iii) the physics parameters are systematically pushed out of distribution, and (iv) how a single model pre-trained on a mixture of different physics problems can be adapted to various downstream applications. We find that-when fine-tuned appropriately-transfer learning can help reach desired accuracy levels with orders of magnitude fewer downstream examples (across different tasks that can even be out-of-distribution) than training from scratch, with consistent behavior across a wide range of downstream examples. We also find that fine-tuning these models yields more performance gains as model size increases, compared to training from scratch on new downstream tasks. These results hold for a broad range of PDE learning tasks. All in all, our results demonstrate the potential of the"pre-train and fine-tune"paradigm for SciML problems, demonstrating a path towards building SciML foundation models. We open-source our code for reproducibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Poseidon: Efficient Foundation Models for PDEsMaximilian Herde, Bogdan Raonic, Tobias Rohner, Roger Käppeli 等NeurIPS 2024 · 被引用 235 次
- DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-TrainingZhongkai Hao, Chang Su, Songming Liu, Julius Berner 等ICML 2024 · 被引用 107 次
- Multiple Physics Pretraining for Spatiotemporal Surrogate ModelsMichael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana 等NeurIPS 2024 · 被引用 97 次
- Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEsMd. Ashiqur Rahman, Robert Joseph George, Mogab Elleithy, Daniel V. Leibovici 等NeurIPS 2024 · 被引用 79 次
- Data-Efficient Operator Learning via Unsupervised Pretraining and In-Context LearningWuyang Chen, Jialin Song, Pu Ren, Shashank Subramanian 等NeurIPS 2024 · 被引用 41 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- Characterizing possible failure modes in physics-informed neural networksAditi S. Krishnapriyan, Amir Gholami, Shandian Zhe, Robert M. Kirby 等NeurIPS 2021 · 被引用 1,421 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Transfer Learning Enhanced DeepONet for Long-Time Prediction of Evolution EquationsWuzhe Xu, Yulong Lu, Li WangAAAI 2023 · 被引用 50 次
相关 Paper
- Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics KnowledgeSiying (Sydney) Ma, Mehrdad Momeni Zadeh, Mauricio Soroco, Wuyang Chen 等ICLR 2026 · 被引用 4 次
- RealPDEBench: A Benchmark for Complex Physical Systems with Real-World DataPeiyan Hu, Haodong Feng, Hongyuan Liu, Tongtong Yan 等ICLR 2026 · 被引用 17 次
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang 等ICML 2024 · 被引用 61 次
- Axial Neural Networks for Dimension-Free Foundation ModelsHyunsu Kim, Jonggeon Park, Joan Bruna, Hongseok Yang 等NeurIPS 2025 · 被引用 1 次
- Features are fate: a theory of transfer learning in high-dimensional regressionJavan Tahir, Surya Ganguli, Grant M. RotskoffICML 2025
