Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
Kieran Didi, Zuobai Zhang, Guoqing Zhou, Danny Reidenbach, Zhonglin Cao, Sooyoung Cha, Tomas Geffner, Christian Dallago, Jian Tang, Michael M. Bronstein, Martin Steinegger, Emine Küçükbenli
Abstract
Protein interaction modeling is central to protein design, which has been transformed by machine learning with applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is cast as either conditional generative modeling or sequence optimization via structure predictors ("hallucination"). We argue that this is a false dichotomy and propose Proteína-Complexa, a novel fully atomistic binder generation method unifying both paradigms. We extend recent flow-based latent protein generation architectures and leverage the domaindomain interactions of monomeric computationally predicted protein structures to construct Teddymer, a new large-scale dataset of synthetic binder-target pairs for pretraining. Combined with high-quality experimental multimers, this enables training a strong base model. We then perform inference-time optimization with this generative prior, unifying the strengths of previously distinct generative and hallucination methods. Proteína-Complexa sets a new state of the art in computational binder design benchmarks: it delivers markedly higher in-silico success rates than existing generative approaches, and our novel test-time optimization strategies greatly outperform previous hallucination methods under normalized compute budgets. We also demonstrate interface hydrogen bond optimization, fold class-guided binder generation, and extensions to small molecule targets and enzyme design tasks, again surpassing prior methods. Code, models and new data will be publicly released. * Core contributor. ♢ Equal advising. † Project lead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- La-Proteina: Atomistic Protein Generation via Partially Latent Flow MatchingTomas Geffner, Kieran Didi, Zhonglin Cao, Danny Reidenbach et al.ICLR 2026 · 57 citations
- Coarse-Grained Boltzmann GeneratorsWeilong Chen, Bojun Zhao, Jan Eckwert, Julija ZavadlavICML 2026
- Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingHengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang et al.ICML 2026
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Proteina: Scaling Flow-based Protein Structure Generative ModelsTomas Geffner, Kieran Didi, Zuobai Zhang, Danny Reidenbach et al.ICLR 2025
- NeuralPLexer3: Accurate Biomolecular Complex Structure Prediction with Flow ModelsJarren Zhuoran Qiao, Feizhi Ding, Thomas Dresselhaus, Mia A. Rosenfeld et al.NeurIPS 2025 · 23 citations
- ProtDBench: A Unified Benchmark of Protein Binder Design and EvaluationCong Liu, Milong Ren, Jiaqi Guan, Chengyue Gong et al.ICML 2026
- All-atom inverse protein folding through discrete flow matchingKai Yi, Kiarash Jamali, Sjors H. W. ScheresICML 2025
- PPDiff: Diffusing in Hybrid Sequence-Structure Space for Protein-Protein Complex DesignZhenqiao Song, Tianxiao Li, Lei Li, Martin Renqiang MinICML 2025
