Latent Diffusion Pretraining for Crystal Property Prediction
Shrimon Mukherjee, KISHALAY DAS, Partha Basuchowdhuri, Pawan Goyal, Niloy Ganguly
摘要
Fast and accurate prediction of crystal properties is a central challenge in new materials design. Graph neural networks and Transformer-based models have emerged as powerful tools for this task due to their ability to encode the local structural environment of atoms within a crystal. However, these models are data-hungry and in practice labeled data for crystal properties are very scarce. Pretraining–finetuning strategies, particularly those based on diffusion models, have shown promise in addressing these limitations. In this work, we introduce a novel latent-diffusion based pretraining framework CrysLDNet, designed to mitigate the data scarcity. Our approach integrates a Variational Autoencoder (VAE) with a diffusion model during the pretraining stage. The VAE encoder maps 3D crystal structures into a smooth latent space, within which the diffusion process is applied. This latent diffusion pretraining enables the graph encoder to effectively capture structural and chemical semantics from large-scale unlabeled data, which can then be finetuned for specific property prediction tasks. Comprehensive experiments on popular DFT datasets for property prediction reveal that CrysLDNet significantly outperforms both training-from-scratch and pretrained baselines, with improvements of 4.26% and 4.90% on the JARVIS and MP datasets. Additionally, the learned representations remain robust in sparse-data conditions and are expressive enough to correct DFT errors when finetuned with limited experimental data. Code is available at https://github.com/shrimonmuke0202/CrysLDNet.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- A Diffusion-Based Pre-training Framework for Crystal Property PredictionZixing Song, Ziqiao Meng, Irwin KingAAAI 2024 · 被引用 25 次
- CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline MaterialsKishalay Das, Bidisha Samanta, Pawan Goyal, Seung-Cheol Lee 等AAAI 2023 · 被引用 27 次
- Beyond Structure: Invariant Crystal Property Prediction with Pseudo-Particle Ray DiffractionBin Cao, Yang Liu, Longhan Zhang, Yifan Wu 等ICLR 2026 · 被引用 4 次
- Crystalformer: Infinitely Connected Attention for Periodic Structure EncodingTatsunori Taniai, Ryo Igarashi, Yuta Suzuki, Naoya Chiba 等ICLR 2024 · 被引用 21 次
- 3D Infomax improves GNNs for Molecular Property PredictionHannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou 等ICML 2022 · 被引用 269 次
