PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Andy Xu, Rohan Desai, Larry Wang, Ethan Ritz, Gabriel Hope
Abstract
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not necessarily to produce the correct answer, but instead to produce a diverse array of candidates which satisfy a set of constraints. We study this challenge in the context of materials generation. To this end, we introduce PLaID++, an LLM post-trained for stable and property-guided crystal generation. We find that applying naive preference optimization to a coordinate-based crystal representation leads to mode collapse. Hence, we introduce a compact, symmetry-informed Wyckoff text representation which improves computational efficiency and encourages generalization from physical priors. By encoding symmetry constraints directly into text and guiding model outputs towards desirable chemical space, PLaID++ generates structures that are thermodynamically stable, unique, and novel at a 50% greater rate than prior methods. We further demonstrate that unified training across conditional and unconditional tasks are mutually beneficial in data-sparse regimes. Our work demonstrates the potential of adapting post-training techniques from natural language processing to materials design, paving the way for targeted and efficient discovery of novel materials.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 970fd626-51e4-458f-9089-dd726b49ced6Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Crystal Diffusion Variational Autoencoder for Periodic Material GenerationTian Xie, Xiang Fu, Octavian-Eugen Ganea, Regina Barzilay et al.ICLR 2022 · 394 citations
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsYi-Lun Liao, Brandon M. Wood, Abhishek Das, Tess E. SmidtICLR 2024 · 311 citations
- Crystal Structure Prediction by Joint Equivariant DiffusionRui Jiao, Wenbing Huang, Peijia Lin, Jiaqi Han et al.NeurIPS 2023 · 245 citations
Related papers
- Fine-Tuned Language Models Generate Stable Inorganic Materials as TextNate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson et al.ICLR 2024 · 120 citations
- CrystalICL: Enabling In-Context Learning for Crystal GenerationRuobing Wang, Qiaoyu Tan, Yili Wang, Ying Wang et al.EMNLP 2025 · 3 citations
- Open Materials Generation with Inference-Time Reinforcement LearningPhilipp Höllmer, Stefano MartinianiICML 2026 · 3 citations
- Wyckoff Transformer: Generation of Symmetric CrystalsNikita Kazeev, Wei Nong, Ignat Romanov, Ruiming Zhu et al.ICML 2025
- LLM Meets Diffusion: A Hybrid Framework for Crystal Material GenerationSubhojyoti Khastagir, Kishalay Das, Pawan Goyal, Seung-Cheol Lee et al.NeurIPS 2025 · 14 citations
