Protein Inverse Folding From Structure Feedback
Junde Xu, Zijun Gao, Xinyi Zhou, Jie Hu, Xingyi Cheng, Le Song, Guangyong Chen, Pheng-Ann Heng, Jiezhong Qiu
Abstract
The inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications. Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model. Given a target protein structure, we begin by sampling candidate sequences from the inverse-folding model, then predict the three-dimensional structure of each sequence with the folding model to generate pairwise structuralpreference labels. These labels are used to fine-tune the inverse-folding model under the DPO objective. Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity. Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5% with regard to the baseline model. This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization † .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 135af703-829f-4f56-a086-88ea6f7d8d60Cited by top-tier papers2
- Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak SupervisionJosh Sun, Morteza Babaie, Wenyang hou, Mark Crowley et al.ICML 2026
- FIDIA: Function-Informed Sequence Design via Inference-Aligned Policy OptimizationMinghan Li, fengji Li, Yilin Tao, Yue DengICML 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend et al.ICLR 2021 · 627 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
Related papers
- PRISM: Enhancing PRotein Inverse Folding through Fine- Grained Retrieval on Structure-Sequence Multimodal RepresentationsSazan Mahbub, Souvik Kundu, Eric P. XingICLR 2026 · 6 citations
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentJunqi Liu, Xiaoyang Hou, Chence Shi, Xin Liu et al.ICLR 2026 · 5 citations
- KW-Design: Pushing the Limit of Protein Design via Knowledge RefinementZhangyang Gao, Cheng Tan, Xingran Chen, Yijie Zhang et al.ICLR 2024 · 21 citations
- SIPF: Sampling Method for Inverse Protein FoldingTianfan Fu, Jimeng SunKDD 2022 · 3 citations
- DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein DesignYanting Li, Zikang Wang, Jiyue Jiang, Ziqian Lin et al.AAAI 2026
