Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
Kieran Didi, Zuobai Zhang, Guoqing Zhou, Danny Reidenbach, Zhonglin Cao, Sooyoung Cha, Tomas Geffner, Christian Dallago, Jian Tang, Michael M. Bronstein, Martin Steinegger, Emine Küçükbenli
摘要
Protein interaction modeling is central to protein design, which has been transformed by machine learning with applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is cast as either conditional generative modeling or sequence optimization via structure predictors ("hallucination"). We argue that this is a false dichotomy and propose Proteína-Complexa, a novel fully atomistic binder generation method unifying both paradigms. We extend recent flow-based latent protein generation architectures and leverage the domaindomain interactions of monomeric computationally predicted protein structures to construct Teddymer, a new large-scale dataset of synthetic binder-target pairs for pretraining. Combined with high-quality experimental multimers, this enables training a strong base model. We then perform inference-time optimization with this generative prior, unifying the strengths of previously distinct generative and hallucination methods. Proteína-Complexa sets a new state of the art in computational binder design benchmarks: it delivers markedly higher in-silico success rates than existing generative approaches, and our novel test-time optimization strategies greatly outperform previous hallucination methods under normalized compute budgets. We also demonstrate interface hydrogen bond optimization, fold class-guided binder generation, and extensions to small molecule targets and enzyme design tasks, again surpassing prior methods. Code, models and new data will be publicly released. * Core contributor. ♢ Equal advising. † Project lead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- La-Proteina: Atomistic Protein Generation via Partially Latent Flow MatchingTomas Geffner, Kieran Didi, Zhonglin Cao, Danny Reidenbach 等ICLR 2026 · 被引用 57 次
- Coarse-Grained Boltzmann GeneratorsWeilong Chen, Bojun Zhao, Jan Eckwert, Julija ZavadlavICML 2026
- Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingHengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang 等ICML 2026
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Proteina: Scaling Flow-based Protein Structure Generative ModelsTomas Geffner, Kieran Didi, Zuobai Zhang, Danny Reidenbach 等ICLR 2025
- NeuralPLexer3: Accurate Biomolecular Complex Structure Prediction with Flow ModelsJarren Zhuoran Qiao, Feizhi Ding, Thomas Dresselhaus, Mia A. Rosenfeld 等NeurIPS 2025 · 被引用 23 次
- ProtDBench: A Unified Benchmark of Protein Binder Design and EvaluationCong Liu, Milong Ren, Jiaqi Guan, Chengyue Gong 等ICML 2026
- All-atom inverse protein folding through discrete flow matchingKai Yi, Kiarash Jamali, Sjors H. W. ScheresICML 2025
- PPDiff: Diffusing in Hybrid Sequence-Structure Space for Protein-Protein Complex DesignZhenqiao Song, Tianxiao Li, Lei Li, Martin Renqiang MinICML 2025
