Data Distillation for extrapolative protein design through exact preference optimization
Mostafa Karimi, Sharmi Banerjee, Tommi S. Jaakkola, Bella Dubrov, Shang Shang, Ron Benson
摘要
The goal of protein design typically involves increasing fitness (extrapolating) beyond what is seen during training (e.g., towards higher stability, stronger binding affinity, etc.). State-of-the-art methods assume that one can safely steer proteins towards such extrapolated regions by learning from pairs alone. We hypothesize that noisy training pairs are not sufficiently informative to capture the fitness gradient and that models learned from pairs specifically may fail to capture threeway relations important for search, e.g., how two alternatives fair relative to a seed. Building on the success of preference alignment models in large language models, we introduce a progressive search method for extrapolative protein design by directly distilling into the model relevant triplet relations. We evaluated our model's performance in designing AAV and GFP proteins and demonstrated that the proposed framework significantly improves effectiveness in extrapolation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Advancing Protein Design via Multi-Agent Reinforcement Learning with Pareto-Based Collaborative OptimizationMingming Zhu, Jiahua Rao, Xiaoyu Chen, Qianmu Yuan 等AAAI 2026 · 被引用 1 次
- Generative property enhancer: implicit guided generation through conditional density estimationPedro O. Pinheiro, Pan Kessel, Aya Abdelsalam Ismail, Sai Pooja Mahajan 等NeurIPS 2025
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du 等ICLR 2021 · 被引用 364 次
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho 等NeurIPS 2024 · 被引用 287 次
相关 Paper
- Protriever: End-to-End Differentiable Protein Homology Search for Fitness PredictionRuben Weitzman, Peter Mørch Groth, Lood van Niekerk, Aoi Otani 等ICML 2025
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsMichal Kmicikiewicz, Vincent Fortuin, Ewa SzczurekNeurIPS 2025 · 被引用 4 次
- Extrapolative Controlled Sequence Generation via Iterative RefinementVishakh Padmakumar, Richard Yuanzhe Pang, He He, Ankur P. ParikhICML 2023 · 被引用 13 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentJunqi Liu, Xiaoyang Hou, Chence Shi, Xin Liu 等ICLR 2026 · 被引用 5 次
