Adaptive Protein Tokenization
Rohit Dilip, Ayush Varshney, David Van Valen
摘要
Tokenization is a promising path to multi-modal models capable of jointly understanding protein sequences, structure, and function. Existing protein structure tokenizers create tokens by pooling information from local neighborhoods, an approach that limits their performance on generative and representation tasks. In this work, we present a method for global tokenization of protein structures in which successive tokens contribute increasing levels of detail to a global representation. This change resolves several issues with generative models based on local protein tokenization: it mitigates error accumulation, provides embeddings without sequence-reduction operations, and allows task-specific adaptation of a tokenized sequence's information content. We validate our method on reconstruction, generative, and representation tasks and demonstrate that it matches or outperforms existing models based on local protein structure tokenizers. We show that our adaptive approach enables inference criteria based on the information content of the generated proteins. We validate representations generated from our tokenizer on CATH classification tasks and demonstrate that non-linear probing on our tokenized sequences outperforms equivalent probing on representations from other tokenizers. Finally, we demonstrate how our method supports zero-shot protein shrinking and affinity maturation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- SE(3) diffusion model with application to protein backbone generationJason Yim, Brian L. Trippe, Valentin De Bortoli, Emile Mathieu 等ICML 2023 · 被引用 313 次
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 被引用 276 次
- Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion ModelsTuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine 等NeurIPS 2024 · 被引用 270 次
相关 Paper
- Flow Autoencoders are Effective Protein TokenizersRohit Dilip, Evan Zhang, Ayush Varshney, David Van ValenICLR 2026 · 被引用 5 次
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICLR 2025
- BiHiTo: Biomolecular Hierarchy-inspired TokenizationRuochong Zheng, Yutian Liu, Yian Zhao, Zhiwei Nie 等AAAI 2026
