Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians
Ishan Amin, Sanjeev Raja, Aditi S. Krishnapriyan
Abstract
The foundation model (FM) paradigm is transforming Machine Learning Force Fields (MLFFs), leveraging general-purpose representations and scalable training to perform a variety of computational chemistry tasks. Although MLFF FMs have begun to close the accuracy gap relative to first-principles methods, there is still a strong need for faster inference speed. Additionally, while research is increasingly focused on general-purpose models which transfer across chemical space, practitioners typically only study a small subset of systems at a given time. This underscores the need for fast, specialized MLFFs relevant to specific downstream applications, which preserve test-time physical soundness while maintaining train-time scalability. In this work, we introduce a method for transferring general-purpose representations from MLFF foundation models to smaller, faster MLFFs specialized to specific regions of chemical space. We formulate our approach as a knowledge distillation procedure, where the smaller "student" MLFF is trained to match the Hessians of the energy predictions of the "teacher" foundation model. Our specialized MLFFs can be up to 20 faster than the original foundation model, while retaining, and in some cases exceeding, its performance and that of undistilled models. We also show that distilling from a teacher model with a direct force parameterization into a student model trained with conservative forces (i.e., computed as derivatives of the potential energy) successfully leverages the representations from the large-scale teacher for improved accuracy, while maintaining energy conservation during test-time molecular dynamics simulations. More broadly, our work suggests a new paradigm for MLFF development, in which foundation models are released along with smaller, specialized simulation "engines" for common chemical subsets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42bfd98a-4385-4d4e-a885-d931209b01fbCited by top-tier papers5
- DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic PotentialsKevin Han, Bowen Deng, Amir Barati Farimani, Gerbrand CederICLR 2026 · 10 citations
- Learning Smooth and Expressive Interatomic Potentials for Physical Property PredictionXiang Fu, Brandon M. Wood, Luis Barroso-Luque, Daniel S. Levine et al.ICML 2025
- Speculative Sampling For Faster Molecular DynamicsArthur Kosmala, Stephan Günnemann, Meng Gao, Brandon WoodICML 2026
- PFT: Phonon Fine-tuning for Machine Learned Interatomic PotentialsTeddy Koker, Abhijeet Gangan, Mit Kotak, Jaime Marian et al.ICML 2026
- Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learned Force FieldsYunyang Li, Lin Huang, Luojia Xia, Wenhe Zhang et al.ICML 2026
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- GemNet: Universal Directional Graph Neural Networks for MoleculesJohannes Gasteiger, Florian Becker, Stephan GünnemannNeurIPS 2021 · 665 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Smooth, exact rotational symmetrization for deep learning on point cloudsSergey Pozdnyakov, Michele CeriottiNeurIPS 2023 · 68 citations
Related papers
- Machine Learning Force Fields with Data Cost Aware TrainingAlexander Bukharin, Tianyi Liu, Shengjie Wang, Simiao Zuo et al.ICML 2023 · 1 citation
- Accelerating Molecular Graph Neural Networks via Knowledge DistillationFilip Ekström Kelvinius, Dimitar Georgiev, Artur P. Toshev, Johannes GasteigerNeurIPS 2023 · 22 citations
- The dark side of the forces: assessing non-conservative force models for atomistic machine learningFilippo Bigi, Marcel F. Langer, Michele CeriottiICML 2025
- Foundry: Distilling 3D Foundation Models for the EdgeGuillaume Letellier, Siddharth Srivastava, Frédéric Jurie, Gaurav SharmaCVPR 2026 · 1 citation
- Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer BehaviorShashank Subramanian, Peter Harrington, Kurt Keutzer, Wahid Bhimji et al.NeurIPS 2023 · 173 citations
