Adaptive Protein Tokenization
Rohit Dilip, Ayush Varshney, David Van Valen
Abstract
Tokenization is a promising path to multi-modal models capable of jointly understanding protein sequences, structure, and function. Existing protein structure tokenizers create tokens by pooling information from local neighborhoods, an approach that limits their performance on generative and representation tasks. In this work, we present a method for global tokenization of protein structures in which successive tokens contribute increasing levels of detail to a global representation. This change resolves several issues with generative models based on local protein tokenization: it mitigates error accumulation, provides embeddings without sequence-reduction operations, and allows task-specific adaptation of a tokenized sequence's information content. We validate our method on reconstruction, generative, and representation tasks and demonstrate that it matches or outperforms existing models based on local protein structure tokenizers. We show that our adaptive approach enables inference criteria based on the information content of the generated proteins. We validate representations generated from our tokenizer on CATH classification tasks and demonstrate that non-linear probing on our tokenized sequences outperforms equivalent probing on representations from other tokenizers. Finally, we demonstrate how our method supports zero-shot protein shrinking and affinity maturation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09c4c735-8ec9-4dcb-ae16-0c792e650247Builds on21
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- SE(3) diffusion model with application to protein backbone generationJason Yim, Brian L. Trippe, Valentin De Bortoli, Emile Mathieu et al.ICML 2023 · 313 citations
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 276 citations
- Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion ModelsTuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine et al.NeurIPS 2024 · 270 citations
Related papers
- Flow Autoencoders are Effective Protein TokenizersRohit Dilip, Evan Zhang, Ayush Varshney, David Van ValenICLR 2026 · 5 citations
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong et al.NeurIPS 2024 · 96 citations
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICLR 2025
- BiHiTo: Biomolecular Hierarchy-inspired TokenizationRuochong Zheng, Yutian Liu, Yian Zhao, Zhiwei Nie et al.AAAI 2026
