LICO: Large Language Models for In-Context Molecular Optimization
Tung Nguyen, Aditya Grover
Abstract
Optimizing black-box functions is a fundamental problem in science and engineering. To solve this problem, many approaches learn a surrogate function that estimates the underlying objective from limited historical evaluations. Large Language Models (LLMs), with their strong pattern-matching capabilities via pretraining on vast amounts of data, stand out as a potential candidate for surrogate modeling. However, directly prompting a pretrained language model to produce predictions is not feasible in many scientific domains due to the scarcity of domain-specific data in the pretraining corpora and the challenges of articulating complex problems in natural language. In this work, we introduce LICO, a general-purpose model that extends arbitrary base LLMs for black-box optimization, with a particular application to the molecular domain. To achieve this, we equip the language model with a separate embedding layer and prediction layer, and train the model to perform in-context predictions on a diverse set of functions defined over the domain. Once trained, LICO can generalize to unseen molecule properties simply via in-context prompting. LICO performs competitively on PMO, a challenging molecular optimization benchmark comprising 23 objective functions, and achieves state-of-the-art performance on its low-budget version PMO-1K.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe000dc7-1de5-4c58-a073-ef289333ed8eCited by top-tier papers6
- Probing the Decision Boundaries of In-context Learning in Large Language ModelsSiyan Zhao, Tung Nguyen, Aditya GroverNeurIPS 2024 · 26 citations
- Pretrained Optimization Model for Zero-Shot Black Box OptimizationXiaobin Li, Kai Wu, Yujian Betterest Li, Xiaoyu Zhang et al.NeurIPS 2024 · 23 citations
- Enhancing Zero-Shot Black-Box Optimization via Pretrained Models with Efficient Population Modeling, Interaction, and Stable Gradient ApproximationMuqi Han, Xiaobin Li, Kai Wu, Xiaoyu Zhang et al.NeurIPS 2025 · 9 citations
- Reference-guided Policy Optimization for Molecular Optimization via LLM ReasoningXuan Li, Zhanke Zhou, Zongze Li, Jiangchao Yao et al.ICLR 2026 · 5 citations
- Assay2Mol: Large Language Model-based Drug Design Using BioAssay ContextYifan Deng, Spencer S. Ericksen, Anthony GitterEMNLP 2025 · 1 citation
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
Related papers
- A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?Agustinus Kristiadi, Felix Strieth-Kalthoff, Marta Skreta, Pascal Poupart et al.ICML 2024 · 55 citations
- A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen et al.ACL 2026 · 10 citations
- Improving LLM-based Global Optimization with Search Space PartitioningAndrej Schwanke, Lyubomir Ivanov, David Salinas, Fabio Ferreira et al.ICLR 2026 · 7 citations
- Tag-LLM: Repurposing General-Purpose LLMs for Specialized DomainsJunhong Shen, Neil A. Tenenholtz, James Brian Hall, David Alvarez-Melis et al.ICML 2024 · 60 citations
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 33 citations
