A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
Agustinus Kristiadi, Felix Strieth-Kalthoff, Marta Skreta, Pascal Poupart, Alán Aspuru-Guzik, Geoff Pleiss
摘要
Automation is one of the cornerstones of contemporary material discovery. Bayesian optimization (BO) is an essential part of such workflows, enabling scientists to leverage prior domain knowledge into efficient exploration of a large molecular space. While such prior knowledge can take many forms, there has been significant fanfare around the ancillary scientific knowledge encapsulated in large language models (LLMs). However, existing work thus far has only explored LLMs for heuristic materials searches. Indeed, recent work obtains the uncertainty estimate -- an integral part of BO -- from point-estimated, non-Bayesian LLMs. In this work, we study the question of whether LLMs are actually useful to accelerate principled Bayesian optimization in the molecular space. We take a sober, dispassionate stance in answering this question. This is done by carefully (i) viewing LLMs as fixed feature extractors for standard but principled BO surrogate models and by (ii) leveraging parameter-efficient finetuning methods and Bayesian neural networks to obtain the posterior of the LLM surrogate. Our extensive experiments with real-world chemistry problems show that LLMs can be useful for BO over molecules, but only if they have been pretrained or finetuned with domain-specific data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Efficient Evolutionary Search Over Chemical Space with Large Language ModelsHaorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao 等ICLR 2025
- Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical SpaceMohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay 等ICML 2026
- Searching for Optimal Solutions with LLMs via Bayesian OptimizationDhruv Agarwal, Manoj Ghuhan Arivazhagan, Rajarshi Das, Sandesh Swamy 等ICLR 2025
- Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent InferenceXuanning Hu, Anchen Li, Qianli Xing, Jinglong Ji 等WWW 2026
- LICO: Large Language Models for In-Context Molecular OptimizationTung Nguyen, Aditya GroverICLR 2025
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
相关 Paper
- LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective ExperimentationZhuo Chen, Xinzhe Yuan, Jianshu Zhang, Jinzong Dong 等ICML 2026 · 被引用 2 次
- Unleashing LLMs in Bayesian Optimization: Preference-Guided Framework for Scientific DiscoveryXinzhe Yuan, Zhuo Chen, Jianshu Zhang, Huan Xiong 等ICLR 2026 · 被引用 7 次
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 被引用 143 次
- Scaling Multi-Task Bayesian Optimization with Large Language ModelsYimeng Zeng, Natalie Maus, Haydn Thomas Jones, Jeffrey Tao 等ICLR 2026 · 被引用 1 次
- FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearchVirginia Aglietti, Ira Ktena, Jessica Schrouff, Eleni Sgouritsa 等ICML 2025
