Model-based reinforcement learning for biological sequence design
Christof Angermüller, David Dohan, David Belanger, Ramya Deshpande, Kevin Murphy, Lucy J. Colwell
Abstract
The ability to design biological structures such as DNA or proteins would have considerable medical and industrial impact. Doing so presents a challenging black-box optimization problem characterized by the large-batch, low round setting due to the need for labor-intensive wet lab evaluations. In response, we propose using reinforcement learning (RL) based on proximal-policy optimization (PPO) for biological sequence design. RL provides a flexible framework for optimization generative sequence models to achieve specific criteria, such as diversity among the high-quality sequences discovered. We propose a model-based variant of PPO, DyNA-PPO, to improve sample efficiency, where the policy for a new round is trained offline using a simulator fit on functional measurements from prior rounds. To accommodate the growing number of observations across rounds, the simulator model is automatically selected at each round from a pool of diverse models of varying capacity. On the tasks of designing DNA transcription factor binding sites, designing antimicrobial proteins, and optimizing the energy of Ising models based on protein structure, we find that DyNA-PPO performs significantly better than existing methods in settings in which modeling is feasible, while still not performing worse in situations in which a reliable model cannot be learned.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fa128c57-e5cd-455e-8fac-71147b2fe99dCited by top-tier papers62
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- Biological Sequence Design with GFlowNetsMoksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks et al.ICML 2022 · 224 citations
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Learning Collaborative Policies to Solve NP-hard Routing ProblemsMinsu Kim, Jinkyoo Park, Joungho KimNeurIPS 2021 · 175 citations
Related papers
- Designing Biological Sequences without Prior Knowledge Using Evolutionary Reinforcement LearningXi Zeng, Xiaotian Hao, Hongyao Tang, Zhentao Tang et al.AAAI 2024 · 2 citations
- Aligning Transformers with Continuous Feedback via Energy Rank AlignmentShriram Chennakesavalu, Frank Hu, Sebastian Ibarraran, Grant M. RotskoffNeurIPS 2025 · 7 citations
- Bridging Model-Based Optimization and Generative Modeling via Conservative Fine-Tuning of Diffusion ModelsMasatoshi Uehara, Yulai Zhao, Ehsan Hajiramezanali, Gabriele Scalia et al.NeurIPS 2024 · 31 citations
- Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RLXingyu Chen, Shihao Ma, Runsheng Lin, Jiecong Lin et al.NeurIPS 2025 · 5 citations
- Improved Off-policy Reinforcement Learning in Biological Sequence DesignHyeonah Kim, Minsu Kim, Taeyoung Yun, Sanghyeok Choi et al.ICML 2025
