Offline Model-based Optimization for Real-World Molecular Discovery
Dong-Hee Shin, Young-Han Son, Hyun Jung Lee, Deok-Joong Lee, Tae-Eui Kam
Abstract
Molecular discovery has attracted significant attention in scientific fields for its ability to generate novel molecules with desirable properties. Although numerous methods have been developed to tackle this problem, most rely on an online setting that requires repeated online evaluation of candidate molecules using the oracle. However, in real-world molecular discovery, the oracle is often represented by wet lab experiments, making this online setting impractical due to the significant time and resource demands. To fill this gap, we propose the Molecular Stitching (MolStitch) framework, which utilizes a fixed offline dataset to explore and optimize molecules without the need for repeated oracle evaluations. Specifically, Mol-Stitch leverages existing molecules from the offline dataset to generate novel 'stitched molecules' that combine their desirable properties. These stitched molecules are then used as training samples to fine-tune the generative model using preference optimization techniques. Experimental results on various offline multi-objective molecular optimization problems validate the effectiveness of MolStitch. The source code is available online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d4a0a4d-1ba7-4f14-853e-ee563a41248dCited by top-tier papers1
Ask how each one uses itBuilds on33
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 352 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Statistical Rejection Sampling Improves Preference OptimizationTianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman et al.ICLR 2024 · 346 citations
Related papers
- Molecule Generation with Fragment Retrieval AugmentationSeul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu et al.NeurIPS 2024 · 36 citations
- MARS: Markov Molecular Sampling for Multi-objective Drug DiscoveryYutong Xie, Chence Shi, Hao Zhou, Yuwei Yang et al.ICLR 2021 · 186 citations
- Refine Drugs, Don’t Complete Them: Uniform-Source Discrete Flows for Fragment-Based Drug DiscoveryBenno Kaech, Luis Wyss, Karsten Borgwardt, Gianvito GrassoICLR 2026 · 3 citations
- Training-free Multi-objective Diffusion Model for 3D Molecule GenerationXu Han, Caihua Shan, Yifei Shen, Can Xu et al.ICLR 2024 · 20 citations
- GenMol: A Drug Discovery Generalist with Discrete DiffusionSeul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu et al.ICML 2025
