De novo Protein Design Using Geometric Vector Field Networks
Weian Mao, Muzhi Zhu, Zheng Sun, Shuaike Shen, Lin Yuanbo Wu, Hao Chen, Chunhua Shen
Abstract
Innovations like protein diffusion have enabled significant progress in de novo protein design, which is a vital topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise features, such as angles and distances between atoms, which are not available in this context. Thus far, only several simple encoders, such as IPA (Jumper et al., 2021) , have been proposed for this scenario, exposing the frame modeling as a bottleneck. In this work, we proffer the Vector Field Network (VFN), which enables network layers to perform learnable vector computations between coordinates of frame-anchored virtual atoms, thus achieving a higher capability for modeling frames. The vector computation operates in a manner similar to a linear layer, with each input channel receiving 3D virtual atom coordinates instead of scalar values. The multiple feature vectors output by the vector computation are then used to update the residue representations and virtual atom coordinates via attention aggregation. Remarkably, VFN also excels in modeling both frames and atoms, as the real atoms can be treated as the virtual atoms for modeling, positioning VFN as a potential universal encoder. In protein diffusion (frame modeling), VFN exhibits an impressive performance advantage over IPA, excelling in terms of both designability (67.04% vs. 53.58%) and diversity (66.54% vs. 51.98%). In inverse folding (frame and atom modeling), VFN outperforms the previous SoTA model, PiFold (54.7% vs. 51.66%), on sequence recovery rate. We also propose a method of equipping VFN with the ESM model (Lin et al., 2023) , which significantly surpasses the previous ESM-based SoTA (62.67% vs. 55.65%), LM-Design (Zheng et al., 2023), by a substantial margin. * WM, ZS and MZ contributed equally. Work was done when WM was visiting Zhejiang University.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6bef03e-cb1c-44ae-8208-4e6aad0eaf0cCited by top-tier papers9
- UniIF: Unified Molecule Inverse FoldingZhangyang Gao, Jue Wang, Cheng Tan, Lirong Wu et al.NeurIPS 2024 · 17 citations
- Bridge-IF: Learning Inverse Protein Folding with Markov BridgesYiheng Zhu, Jialu Wu, Qiuyi Li, Jiahuan Yan et al.NeurIPS 2024 · 16 citations
- ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree SearchMengdi Liu, Xiaoxue Cheng, Zhangyang Gao, Hong Chang et al.NeurIPS 2025 · 10 citations
- PDFBench: A Benchmark for De Novo Protein Design from FunctionJiahao Kuang, Nuowei Liu, Changzhi Sun, Jie Wang et al.ICML 2026 · 10 citations
- PRISM: Enhancing PRotein Inverse Folding through Fine- Grained Retrieval on Structure-Sequence Multimodal RepresentationsSazan Mahbub, Souvik Kundu, Eric P. XingICLR 2026 · 6 citations
Builds on9
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend et al.ICLR 2021 · 627 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- SE(3) diffusion model with application to protein backbone generationJason Yim, Brian L. Trippe, Valentin De Bortoli, Emile Mathieu et al.ICML 2023 · 313 citations
- ComENet: Towards Complete and Efficient Message Passing for 3D Molecular GraphsLimei Wang, Yi Liu, Yuchao Lin, Haoran Liu et al.NeurIPS 2022 · 130 citations
Related papers
- Generating Novel, Designable, and Diverse Protein Structures by Equivariantly Diffusing Oriented Residue CloudsYeqing Lin, Mohammed AlQuraishiICML 2023 · 105 citations
- EVA: Geometric Inverse Design for Fast Protein Motif-Scaffolding with Coupled FlowYufei Huang, Yunshu Liu, Lirong Wu, Haitao Lin et al.ICLR 2025
- CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based ModelMilong Ren, Tian Zhu, Haicang ZhangICML 2024 · 13 citations
- Sequence-Augmented SE(3)-Flow Matching For Conditional Protein GenerationGuillaume Huguet, James Vuckovic, Kilian Fatras, Eric Thibodeau-Laufer et al.NeurIPS 2024 · 32 citations
- DecompDiff: Diffusion Models with Decomposed Priors for Structure-Based Drug DesignJiaqi Guan, Xiangxin Zhou, Yuwei Yang, Yu Bao et al.ICML 2023 · 115 citations
