Data Distillation for extrapolative protein design through exact preference optimization
Mostafa Karimi, Sharmi Banerjee, Tommi S. Jaakkola, Bella Dubrov, Shang Shang, Ron Benson
Abstract
The goal of protein design typically involves increasing fitness (extrapolating) beyond what is seen during training (e.g., towards higher stability, stronger binding affinity, etc.). State-of-the-art methods assume that one can safely steer proteins towards such extrapolated regions by learning from pairs alone. We hypothesize that noisy training pairs are not sufficiently informative to capture the fitness gradient and that models learned from pairs specifically may fail to capture threeway relations important for search, e.g., how two alternatives fair relative to a seed. Building on the success of preference alignment models in large language models, we introduce a progressive search method for extrapolative protein design by directly distilling into the model relevant triplet relations. We evaluated our model's performance in designing AAV and GFP proteins and demonstrated that the proposed framework significantly improves effectiveness in extrapolation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d25095d9-e9e1-43a0-89e8-b49051b40366Cited by top-tier papers2
- Advancing Protein Design via Multi-Agent Reinforcement Learning with Pareto-Based Collaborative OptimizationMingming Zhu, Jiahua Rao, Xiaoyu Chen, Qianmu Yuan et al.AAAI 2026 · 1 citation
- Generative property enhancer: implicit guided generation through conditional density estimationPedro O. Pinheiro, Pan Kessel, Aya Abdelsalam Ismail, Sai Pooja Mahajan et al.NeurIPS 2025
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du et al.ICLR 2021 · 364 citations
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho et al.NeurIPS 2024 · 287 citations
Related papers
- Protriever: End-to-End Differentiable Protein Homology Search for Fitness PredictionRuben Weitzman, Peter Mørch Groth, Lood van Niekerk, Aoi Otani et al.ICML 2025
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsMichal Kmicikiewicz, Vincent Fortuin, Ewa SzczurekNeurIPS 2025 · 4 citations
- Extrapolative Controlled Sequence Generation via Iterative RefinementVishakh Padmakumar, Richard Yuanzhe Pang, He He, Ankur P. ParikhICML 2023 · 13 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentJunqi Liu, Xiaoyang Hou, Chence Shi, Xin Liu et al.ICLR 2026 · 5 citations
