Learning Protein Structure-Function Relationships through Knowledge-guided Representation Decomposition
Mingqing Wang, Zhiwei Nie, ATHANASIOS VASILAKOS, Yonghong He, Zhixiang Ren
摘要
Proteins encode diverse functions within complex three-dimensional structures, yet most deep learning representations remain highly entangled, obscuring the biophysical signals that underlie function. Here we introduce ProtDiS, a knowledge-guided framework that decomposes pretrained protein micro-environment embeddings into biologically grounded and task-relevant dimensions. Inspired by the information bottleneck principle, ProtDiS learns representations that balance informativeness and compression, yielding structural features that are more specific, independent, and information-efficient, and achieving consistent improvements across twelve downstream tasks, with the largest gains under structure-based splits. Protein- and residue-level analyses further show that ProtDiS differentiates proteins with similar folds but divergent functions and captures fine-grained biophysical signals critical. These findings suggest that knowledge–guided decomposition provides a general and interpretable approach for structuring latent spaces in protein structural modeling. The source code and implementation details are publicly available at https://github.com/AI-HPC-Research-Team/ProtDiS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- Disentangled Information BottleneckZiqi Pan, Li Niu, Jianfu Zhang, Liqing ZhangAAAI 2021 · 被引用 55 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- Continuous-Discrete Convolution for Geometry-Sequence Modeling in ProteinsHehe Fan, Zhangyang Wang, Yi Yang, Mohan S. KankanhalliICLR 2023
- From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language ModelsEtowah Adams, Liam Bai, Minji Lee, Yiyang Yu 等ICML 2025
相关 Paper
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersXiangyu Liu, Haodi Lei, Yi Liu, Yang Liu 等AAAI 2026 · 被引用 2 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
- Multi-level Protein Structure Pre-training via Prompt LearningZeyuan Wang, Qiang Zhang, Shuangwei Hu, Haoran Yu 等ICLR 2023
- Explaining A Black-box By Using A Deep Variational Information Bottleneck ApproachSeo-Jin Bang, Pengtao Xie, Heewook Lee, Wei Wu 等AAAI 2021 · 被引用 33 次
- Protein Circuit Tracing via Cross-layer TranscodersDarin Tsui, Kunal Talreja, Daniel Saeedi, Amirali AghazadehICML 2026 · 被引用 4 次
