Protein Inverse Folding From Structure Feedback
Junde Xu, Zijun Gao, Xinyi Zhou, Jie Hu, Xingyi Cheng, Le Song, Guangyong Chen, Pheng-Ann Heng, Jiezhong Qiu
摘要
The inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications. Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model. Given a target protein structure, we begin by sampling candidate sequences from the inverse-folding model, then predict the three-dimensional structure of each sequence with the folding model to generate pairwise structuralpreference labels. These labels are used to fine-tune the inverse-folding model under the DPO objective. Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity. Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5% with regard to the baseline model. This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization † .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak SupervisionJosh Sun, Morteza Babaie, Wenyang hou, Mark Crowley 等ICML 2026
- FIDIA: Function-Informed Sequence Design via Inference-Aligned Policy OptimizationMinghan Li, fengji Li, Yilin Tao, Yue DengICML 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
相关 Paper
- PRISM: Enhancing PRotein Inverse Folding through Fine- Grained Retrieval on Structure-Sequence Multimodal RepresentationsSazan Mahbub, Souvik Kundu, Eric P. XingICLR 2026 · 被引用 6 次
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentJunqi Liu, Xiaoyang Hou, Chence Shi, Xin Liu 等ICLR 2026 · 被引用 5 次
- KW-Design: Pushing the Limit of Protein Design via Knowledge RefinementZhangyang Gao, Cheng Tan, Xingran Chen, Yijie Zhang 等ICLR 2024 · 被引用 21 次
- SIPF: Sampling Method for Inverse Protein FoldingTianfan Fu, Jimeng SunKDD 2022 · 被引用 3 次
- DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein DesignYanting Li, Zikang Wang, Jiyue Jiang, Ziqian Lin 等AAAI 2026
