GeLLM³O: Generalizing Large Language Models for Multi-property Molecule Optimization
Vishal Dey, Xiao Hu, Xia Ning
Abstract
Despite recent advancements, most computational methods for molecule optimization are constrained to single-or double-property optimization tasks and suffer from poor scalability and generalizability to novel optimization tasks. Meanwhile, Large Language Models (LLMs) demonstrate remarkable out-of-domain generalizability to novel tasks. To demonstrate LLMs' potential for molecule optimization, we introduce MuMOInstruct, the first high-quality instruction-tuning dataset specifically focused on complex multi-property molecule optimization tasks. Leveraging MuMOInstruct, we develop GeLLM 3 Os, a series of instruction-tuned LLMs for molecule optimization. Extensive evaluations across 5 in-domain and 5 outof-domain tasks demonstrate that GeLLM 3 Os consistently outperform state-of-the-art baselines. GeLLM 3 Os also exhibit outstanding zeroshot generalization to unseen tasks, significantly outperforming powerful closed-source LLMs. Such strong generalizability demonstrates the tremendous potential of GeLLM 3 Os as foundational models for molecule optimization, thereby tackling novel optimization tasks without resource-intensive retraining. MuMOInstruct, models, and code are accessible through https://github.com/ninglab/ GeLLMO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- MARS: Markov Molecular Sampling for Multi-objective Drug DiscoveryYutong Xie, Chence Shi, Hao Zhou, Yuwei Yang et al.ICLR 2021 · 186 citations
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu et al.ICLR 2024 · 137 citations
- MIMOSA: Multi-constraint Molecule Sampling for Molecule OptimizationTianfan Fu, Cao Xiao, Xinhao Li, Lucas M. Glass et al.AAAI 2021 · 94 citations
Related papers
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song et al.NeurIPS 2025 · 5 citations
- A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to OptimizationZiqing Wang, Kexin Zhang, Zihan Zhao, Yibo Wen et al.ACL 2026 · 10 citations
- eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction DataBo Peng, Xinyi Ling, Ziru Chen, Huan Sun et al.ICML 2024 · 53 citations
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 33 citations
- Towards 3D Molecule-Text Interpretation in Language ModelsSihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang et al.ICLR 2024 · 87 citations
