Property-Driven Protein Inverse Folding with Multi-Objective Preference Alignment
Junqi Liu, Xiaoyang Hou, Chence Shi, Xin Liu, Zhi Yang, Jian Tang
Abstract
Protein sequence design must balance designability, defined as the ability to recover a target backbone, with multiple, often competing, developability properties such as solubility, thermostability, and expression. Existing approaches address these properties through post hoc mutation, inference-time biasing, or retraining on property-specific subsets, yet they are target dependent and demand substantial domain expertise or careful hyperparameter tuning. In this paper, we introduce Pro-tAlign, a multi-objective preference alignment framework that fine-tunes pretrained inverse folding models to satisfy diverse developability objectives while preserving structural fidelity. ProtAlign employs a semi-online Direct Preference Optimization strategy with a flexible preference margin to mitigate conflicts among competing objectives and constructs preference pairs using in silico property predictors. Applied to the widely used ProteinMPNN backbone, the resulting model MoMPNN enhances developability without compromising designability across tasks including sequence design for CATH 4.3 crystal structures, de novo generated backbones, and real-world binder design scenarios, making it an appealing framework for practical protein sequence design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 119eca35-6401-4d5e-927d-a667e10f5872Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend et al.ICLR 2021 · 627 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
Related papers
- DualMPNN: Harnessing Structural Alignments for High-Recovery Inverse Protein FoldingXuhui Liao, Qiyu Wang, Zhiqiang Liang, Liwei Xiao et al.NeurIPS 2025 · 2 citations
- Multi-state Protein Sequence Design with DynamicMPNNAlex Abrudan, Sebastian Pujalte Ojeda, Chaitanya K. Joshi, Matthew Greenig et al.ICLR 2026 · 5 citations
- Protein Inverse Folding From Structure FeedbackJunde Xu, Zijun Gao, Xinyi Zhou, Jie Hu et al.NeurIPS 2025 · 9 citations
- Advancing Protein Design via Multi-Agent Reinforcement Learning with Pareto-Based Collaborative OptimizationMingming Zhu, Jiahua Rao, Xiaoyu Chen, Qianmu Yuan et al.AAAI 2026 · 1 citation
- Discrete Diffusion Trajectory Alignment via Stepwise DecompositionJiaqi Han, Austin Wang, Minkai Xu, Wenda Chu et al.ICLR 2026 · 11 citations
