Distilling Structural Representations into Protein Sequence Models
Jeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl, Adam R. Klivans, Daniel Jesus Diaz
摘要
Protein language models, like the popular ESM2, are widely used tools for extracting evolution-based protein representations and have achieved significant success on downstream biological tasks. Representations based on sequence and structure models, however, show significant performance differences depending on the downstream task. A major open problem is to obtain representations that best capture both the evolutionary and structural properties of proteins in general. Here we introduce Implicit Structure Model (ISM), a sequence-only input model with structurally-enriched representations that outperforms state-of-the-art sequence models on several well-studied benchmarks including mutation stability assessment and structure prediction. Our key innovations are a microenvironment-based autoencoder for generating structure tokens and a self-supervised training objective that distills these tokens into ESM2's pre-trained model. We have made ISM's
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Triangle Multiplication is All You Need for Biomolecular Structure RepresentationsJeffrey Ouyang-Zhang, Pranav Murugan, Daniel Jesus Diaz, Gianluca Scarpellini 等ICLR 2026 · 被引用 8 次
- Representing local protein environments with machine learning force fieldsMeital Bojan, Sanketh Vedula, Sai Advaith Maddipatla, Nadav Bojan 等ICLR 2026 · 被引用 4 次
- Greater than the Sum of Its Parts: Building Substructure into Protein Encoding ModelsRobert Calef, Arthur Liang, Manolis Kellis, Marinka ZitnikICLR 2026 · 被引用 2 次
- FlexRibbon: Joint Sequence and Structure Pretraining for Protein ModelingJianwei Zhu, Yu Shi, Ran Bi, Peiran Jin 等ICLR 2026 · 被引用 2 次
- Ambient Proteins - Training Diffusion Models on Noisy StructuresGiannis Daras, Jeffrey Ouyang-Zhang, Krithika Ravishankar, Constantinos Daskalakis 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper6
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan 等ICLR 2024 · 被引用 285 次
- ProSST: Protein Language Modeling with Quantized Structure and Disentangled AttentionMingchen Li, Yang Tan, Xinzhu Ma, Bozitao Zhong 等NeurIPS 2024 · 被引用 96 次
- Exploring evolution-aware & -free protein language models as protein function predictorsMingyang Hu, Fajie Yuan, Kevin Yang, Fusong Ju 等NeurIPS 2022 · 被引用 70 次
- Predicting a Protein's Stability under a Million MutationsJeffrey Ouyang-Zhang, Daniel Jesus Diaz, Adam R. Klivans, Philipp KrähenbühlNeurIPS 2023 · 被引用 35 次
相关 Paper
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- Protein Structure Tokenization: Benchmarking and New RecipeXinyu Yuan, Zichen Wang, Marcus D. Collins, Huzefa RangwalaICML 2025
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu 等ICML 2025
- Co-Generative De Novo Functional Protein DesignXinRui Chen, YIZHEN LUO, Siqi Fan, Zaiqing NieICML 2026
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICLR 2025
