Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
Marco Jiralerspong, Esther Derman, Danilo Vucetic, Nikolay Malkin, Bilun Sun, Tianyu Zhang, Pierre-Luc Bacon, Gauthier Gidel
Abstract
A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL) guided by a proxy reward function to enable this filtering. By employing various forms of entropy regularization, these methods aim to learn samplers that generate diverse candidates that are highly rated by the proxy function. In this work, we make two main contributions. First, we show that these methods are liable to generate overly diverse, suboptimal candidates in large search spaces. To address this issue, we introduce a novel unified operator that combines several regularized RL operators into a general framework that better targets peakier sampling distributions. Secondly, we offer a novel, robust RL perspective of this filtering process. The regularization can be interpreted as robustness to a compositional form of uncertainty in the proxy function (i.e., the true evaluation of a candidate differs from the proxy's evaluation). Our analysis leads us to a novel, easy-to-use algorithm we name trajectory general mellowmax (TGM): we show it identifies higher quality, diverse candidates than baselines in both synthetic and real-world tasks. Code: https://github.com/marcojira/tgm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcdf1137-2cb4-4b30-8121-7495a8cd6bbcBuilds on13
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Biological Sequence Design with GFlowNetsMoksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks et al.ICML 2022 · 224 citations
- Model-based reinforcement learning for biological sequence designChristof Angermüller, David Dohan, David Belanger, Ramya Deshpande et al.ICLR 2020 · 159 citations
- Learning GFlowNets From Partial Episodes For Improved Convergence And StabilityKanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio et al.ICML 2023 · 138 citations
Related papers
- Feedback Efficient Online Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali et al.ICML 2024 · 47 citations
- Maximum Entropy Reinforcement Learning via Energy-Based Normalizing FlowChen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee et al.NeurIPS 2024 · 29 citations
- Reference-guided Policy Optimization for Molecular Optimization via LLM ReasoningXuan Li, Zhanke Zhou, Zongze Li, Jiangchao Yao et al.ICLR 2026 · 5 citations
- Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement LearningKe Sun, Yingnan Zhao, Enze Shi, Yafei Wang et al.NeurIPS 2025 · 1 citation
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
