Offline Model-Based Optimization via Policy-Guided Gradient Search
Yassine Chemingui, Aryan Deshwal, Trong Nghia Hoang, Janardhan Rao Doppa
Abstract
Offline optimization is an emerging problem in many experimental engineering domains including protein, drug or aircraft design, where online experimentation to collect evaluation data is too expensive or dangerous. To avoid that, one has to optimize an unknown function given only its offline evaluation at a fixed set of inputs. A naive solution to this problem is to learn a surrogate model of the unknown function and optimize this surrogate instead. However, such a naive optimizer is prone to erroneous overestimation of the surrogate (possibly due to over-fitting on a biased sample of function evaluation) on inputs outside the offline dataset. Prior approaches addressing this challenge have primarily focused on learning robust surrogate models. However, their search strategies are derived from the surrogate model rather than the actual offline data. To fill this important gap, we introduce a new learning-to-search perspective for offline optimization by reformulating it as an offline reinforcement learning problem. Our proposed policy-guided gradient search approach explicitly learns the best policy for a given surrogate model created from the offline data. Our empirical results on multiple benchmarks demonstrate that the learned optimization policy can be combined with existing offline surrogates to significantly improve the optimization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99a272bf-705f-4f87-9a23-4bd6d2d29506Cited by top-tier papers16
- Guided Trajectory Generation with Diffusion Models for Offline Model-based OptimizationTaeyoung Yun, Sujin Yun, Jaewoo Lee, Jinkyoo ParkNeurIPS 2024 · 23 citations
- Learning Surrogates for Offline Black-Box Optimization via Gradient MatchingMinh Hoang, Azza Fadhel, Aryan Deshwal, Jana Doppa et al.ICML 2024 · 18 citations
- Offline Multi-Objective OptimizationKe Xue, Rong-Xi Tan, Xiaobin Huang, Chao QianICML 2024 · 14 citations
- Preference-Guided Diffusion for Multi-Objective Offline OptimizationYashas Annadani, Syrine Belakaria, Stefano Ermon, Stefan Bauer et al.NeurIPS 2025 · 12 citations
- Boosting Offline Optimizers with Surrogate SensitivityManh Cuong Dao, Phi Le Nguyen, Truong Thao Nguyen, Trong Nghia HoangICML 2024 · 10 citations
Builds on17
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
Related papers
- Incorporating Surrogate Gradient Norm to Improve Offline Optimization TechniquesCuong Dao, Phi Le Nguyen, Truong Thao Nguyen, Nghia HoangNeurIPS 2024 · 8 citations
- Offline Model-Based Optimization by Learning to RankRong-Xi Tan, Ke Xue, Shen-Huan Lyu, Haopu Shang et al.ICLR 2025
- Generative Adversarial Model-Based Optimization via Source Critic RegularizationMichael S. Yao, Yimeng Zeng, Hamsa Bastani, Jacob R. Gardner et al.NeurIPS 2024 · 14 citations
- Mind the Gap: Offline Policy Optimization for Imperfect RewardsJianxiong Li, Xiao Hu, Haoran Xu, Jingjing Liu et al.ICLR 2023
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
