FoldToken: Learning Protein Language via Vector Quantization and Beyond
Zhangyang Gao, Cheng Tan, Jue Wang, Yufei Huang, Lirong Wu, Stan Z. Li
Abstract
Is there a foreign language describing protein sequences and structures simultaneously? Protein structures, represented by continuous 3D points, have long posed a challenge due to the contrasting modeling paradigms of discrete sequences. We introduce FoldTokenizer to represent protein sequence-structure as discrete symbols. This innovative approach involves projecting residue types and structures into a discrete space, guided by a reconstruction loss for information preservation. We refer to the learned discrete symbols as Fold-Token, and the sequence of FoldTokens serves as a new protein language, transforming the protein sequence-structure into a unified modality. We apply the created protein language on general backbone inpainting and antibody design tasks, building the first GPT-style model (FoldGPT) for sequence-structure co-generation with promising results. Key to our success is the substantial enhancement of the vector quantization module, Soft Conditional Vector Quantization (SoftCVQ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 858aed8e-8573-40e2-ab5f-1ffcf574890cCited by top-tier papers5
- Addressing Representation Collapse in Vector Quantized Models with One Linear LayerYongxin Zhu, Bocheng Li, Yifei Xin, Zhihua Xia et al.ICCV 2025 · 5 citations
- Flow Autoencoders are Effective Protein TokenizersRohit Dilip, Evan Zhang, Ayush Varshney, David Van ValenICLR 2026 · 5 citations
- AlphaFold Database Debiasing for Robust Inverse FoldingCheng Tan, Zhenxiao Cao, Zhangyang Gao, Siyuan Li et al.NeurIPS 2025 · 3 citations
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token AttentionJingbo Zhou, Jun Xia, Siyuan Li, Yunfan Liu et al.WWW 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend et al.ICLR 2021 · 627 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
Related papers
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICLR 2025
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
- Protein Structure Tokenization via Geometric Byte Pair EncodingMichael Sun, Weize Yuan, Gang Liu, Wojciech Matusik et al.ICLR 2026 · 6 citations
- Co-Generative De Novo Functional Protein DesignXinRui Chen, YIZHEN LUO, Siqi Fan, Zaiqing NieICML 2026
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan et al.ICLR 2024 · 285 citations
