Improved Off-policy Reinforcement Learning in Biological Sequence Design
Hyeonah Kim, Minsu Kim, Taeyoung Yun, Sanghyeok Choi, Emmanuel Bengio, Alex Hernández-García, Jinkyoo Park
摘要
Designing biological sequences with desired properties is a significant challenge due to the combinatorially vast search space and the high cost of evaluating each candidate sequence. To address these challenges, reinforcement learning (RL) methods, such as GFlowNets, utilize proxy models for rapid reward evaluation and annotated data for policy training. Although these approaches have shown promise in generating diverse and novel sequences, the limited training data relative to the vast search space often leads to the misspecification of proxy for out-ofdistribution inputs. We introduce δ-Conservative Search, a novel off-policy search method for training GFlowNets designed to improve robustness against proxy misspecification. The key idea is to incorporate conservativeness, controlled by parameter δ, to constrain the search to reliable regions. Specifically, we inject noise into high-score offline sequences by randomly masking tokens with a Bernoulli distribution of parameter δ and then denoise masked tokens using the GFlowNet policy. Additionally, δ is adaptively adjusted based on the uncertainty of the proxy model for each data point. This enables the reflection of proxy uncertainty to determine the level of conservativeness. Experimental results demonstrate that our method consistently outperforms existing machine learning methods in discovering high-score sequences across diverse tasks-including DNA, RNA, protein, and peptide design-especially in large-scale scenarios. The code is available at https://github.com/hyeonahkimm/delta cs .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo 等NeurIPS 2025 · 被引用 13 次
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsMichal Kmicikiewicz, Vincent Fortuin, Ewa SzczurekNeurIPS 2025 · 被引用 4 次
- KeeA*: Epistemic Exploratory A* Search via Knowledge CalibrationDengwei Zhao, Shikui Tu, Yanan Sun, Lei XuNeurIPS 2025
- Revisiting Non-Acyclic GFlowNets in Discrete EnvironmentsNikita Morozov, Ian Maksimov, Daniil Tiapkin, Sergey SamsonovICML 2025
- Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNetsBo Xue, Yunchong Song, Fanghao Shao, Xuekai Zhu 等ICLR 2026
它引用的顶会 Paper15
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Biological Sequence Design with GFlowNetsMoksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks 等ICML 2022 · 被引用 224 次
- Sample-Efficient Optimization in the Latent Space of Deep Generative Models via Weighted RetrainingAustin Tripp, Erik A. Daxberger, José Miguel Hernández-LobatoNeurIPS 2020 · 被引用 186 次
- Model-based reinforcement learning for biological sequence designChristof Angermüller, David Dohan, David Belanger, Ramya Deshpande 等ICLR 2020 · 被引用 159 次
相关 Paper
- Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Ehsan Hajiramezanali, Gabriele Scalia 等NeurIPS 2024 · 被引用 31 次
- Designing Biological Sequences without Prior Knowledge Using Evolutionary Reinforcement LearningXi Zeng, Xiaotian Hao, Hongyao Tang, Zhentao Tang 等AAAI 2024 · 被引用 2 次
- COFlowNet: Conservative Constraints on Flows Enable High-Quality Candidate GenerationYudong Zhang, Xuan Yu, Xu Wang, Zhaoyang Sun 等ICLR 2025
- Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological SequencesMinsu Kim, Federico Berto, Sungsoo Ahn, Jinkyoo ParkNeurIPS 2023 · 被引用 30 次
- Towards Understanding and Improving GFlowNet TrainingMax W. Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas 等ICML 2023 · 被引用 81 次
