Fold2Seq: A Joint Sequence(1D)-Fold(3D) Embedding-based Generative Model for Protein Design
Yue Cao, Payel Das, Vijil Chenthamarakshan, Pin-Yu Chen, Igor Melnyk, Yang Shen
摘要
Designing novel protein sequences for a desired 3D topological fold is a fundamental yet nontrivial task in protein engineering. Challenges exist due to the complex sequence-fold relationship, as well as the difficulties to capture the diversity of the sequences (therefore structures and functions) within a fold. To overcome these challenges, we propose Fold2Seq, a novel transformer-based generative framework for designing protein sequences conditioned on a specific target fold. To model the complex sequence-structure relationship, Fold2Seq jointly learns a sequence embedding using a transformer and a fold embedding from the density of secondary structural elements in 3D voxels. On test sets with single, high-resolution and complete structure inputs for individual folds, our experiments demonstrate improved or comparable performance of Fold2Seq in terms of speed, coverage, and reliability for sequence design, when compared to existing state-of-the-art methods that include data-driven deep generative models and physics-based RosettaDesign. The unique advantages of fold-based Fold2Seq, in comparison to a structure-based deep model and RosettaDesign, become more evident on three additional real-world challenges originating from low-quality, incomplete, or ambiguous input structures. Source code and data are available at https://github.com/IBM/fold2seq.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Iterative Refinement Graph Neural Network for Antibody Sequence-Structure Co-designWengong Jin, Jeremy Wohlwend, Regina Barzilay, Tommi S. JaakkolaICLR 2022 · 被引用 164 次
- Antibody-Antigen Docking and Design via Hierarchical Structure RefinementWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2022 · 被引用 57 次
- PiFold: Toward effective and efficient protein inverse foldingZhangyang Gao, Cheng Tan, Stan Z. LiICLR 2023 · 被引用 50 次
- Reprogramming Pretrained Language Models for Antibody Sequence InfillingIgor Melnyk, Vijil Chenthamarakshan, Pin-Yu Chen, Payel Das 等ICML 2023 · 被引用 40 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
相关 Paper
- Sequence-Augmented SE(3)-Flow Matching For Conditional Protein GenerationGuillaume Huguet, James Vuckovic, Kilian Fatras, Eric Thibodeau-Laufer 等NeurIPS 2024 · 被引用 32 次
- A Joint Diffusion Model with Pre-Trained Priors for RNA Sequence-Structure Co-DesignXiner Li, Masatoshi Uehara, Xingyu Su, Gabriele Scalia 等ICLR 2026
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
- Protein Sequence and Structure Co-Design with Equivariant TranslationChence Shi, Chuanrui Wang, Jiarui Lu, Bozitao Zhong 等ICLR 2023 · 被引用 13 次
- FoldToken: Learning Protein Language via Vector Quantization and BeyondZhangyang Gao, Cheng Tan, Jue Wang, Yufei Huang 等AAAI 2025 · 被引用 29 次
