MeToken: Uniform Micro-environment Token Boosts Post-Translational Modification Prediction
Cheng Tan, Zhenxiao Cao, Zhangyang Gao, Lirong Wu, Siyuan Li, Yufei Huang, Jun Xia, Bozhen Hu, Stan Z. Li
摘要
Post-translational modifications (PTMs) profoundly expand the complexity and functionality of the proteome, regulating protein attributes and interactions that are crucial for biological processes. Accurately predicting PTM sites and their specific types is therefore essential for elucidating protein function and understanding disease mechanisms. Existing computational approaches predominantly focus on protein sequences to predict PTM sites, driven by the recognition of sequence-dependent motifs. However, these approaches often overlook protein structural contexts. In this work, we first compile a large-scale sequence-structure PTM dataset, which serves as the foundation for fair comparison. We introduce the MeToken model, which tokenizes the micro-environment of each amino acid, integrating both sequence and structural information into unified discrete tokens. This model not only captures the typical sequence motifs associated with PTMs but also leverages the spatial arrangements dictated by protein tertiary structures, thus providing a holistic view of the factors influencing PTM sites. Designed to address the long-tail distribution of PTM types, MeToken employs uniform sub-codebooks that ensure even the rarest PTMs are adequately represented and distinguished. We validate the effectiveness and generalizability of MeToken across multiple datasets, demonstrating its superior performance in accurately identifying PTM types. The results underscore the importance of incorporating structural data and highlight MeToken's potential in facilitating accurate and comprehensive PTM predictions, which could significantly impact proteomics research. The code and datasets are available at https://github.com/A4Bio/MeToken.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AlphaFold Database Debiasing for Robust Inverse FoldingCheng Tan, Zhenxiao Cao, Zhangyang Gao, Siyuan Li 等NeurIPS 2025 · 被引用 3 次
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper18
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- Hierarchical Generation of Molecular Graphs using Structural MotifsWengong Jin, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 被引用 356 次
- Antigen-Specific Antibody Design and Optimization with Diffusion-Based Generative Models for Protein StructuresShitong Luo, Yufeng Su, Xingang Peng, Sheng Wang 等NeurIPS 2022 · 被引用 331 次
相关 Paper
- Improving PTM Site Prediction by Coupling of Multi-Granularity Structure and Multi-Scale Sequence RepresentationZhengyi Li, Menglu Li, Lida Zhu, Wen ZhangAAAI 2024 · 被引用 11 次
- MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein EmbeddingLirong Wu, Yijun Tian, Yufei Huang, Siyuan Li 等ICLR 2024 · 被引用 47 次
- Greater than the Sum of Its Parts: Building Substructure into Protein Encoding ModelsRobert Calef, Arthur Liang, Manolis Kellis, Marinka ZitnikICLR 2026 · 被引用 2 次
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- Elucidating the Design Space of Multimodal Protein Language ModelsCheng-Yen Hsieh, Xinyou Wang, Daiheng Zhang, Dongyu Xue 等ICML 2025
