Retrieval-Augmented Language Model for Knowledge-aware Protein Encoding
Jiasheng Zhang, Delvin Ce Zhang, Shuang Liang, Zhengpin Li, Rex Ying, Jie Shao
摘要
Protein language models often struggle to capture biological functions due to their lack of factual knowledge (e.g., gene descriptions). Existing solutions leverage protein knowledge graphs (PKGs) as auxiliary pre-training objectives, but lack explicit integration of task-oriented knowledge, making them suffer from limited knowledge exploitation and catastrophic forgetting. The root cause is that they fail to align PKGs with task-specific data, forcing their knowledge modeling to adapt to the knowledge-isolated nature of downstream tasks. In this paper, we propose Knowledge-aware retrieval augmented protein language model (Kara), achieving the first task-oriented and explicit integration of PKGs and protein language models. With a knowledge retriever learning to predict linkages between PKG and task proteins, Kara unifies the knowledge integration of the pre-training and fine-tuning stages with a structure-based regularization, mitigating catastrophic forgetting. To ensure task-oriented integration, Kara uses contextualized virtual tokens to extract graph context as task-specific knowledge for new proteins. Experiments show that Kara outperforms existing knowledge-enhanced models in 6 representative tasks, achieving on average 5.1% improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question AnsweringXiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla 等NeurIPS 2024 · 被引用 384 次
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 被引用 147 次
- Making Large Language Models Perform Better in Knowledge Graph CompletionYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu 等ACM MM 2024 · 被引用 86 次
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu 等AAAI 2022 · 被引用 36 次
相关 Paper
- Protein Representation Learning via Knowledge Enhanced Primary Structure ReasoningHong-Yu Zhou, Yunxiang Fu, Zhicheng Zhang, Cheng Bian 等ICLR 2023
- Enhancing Conversational Recommender Systems with Tree-Structured Knowledge and Pretrained Language ModelsYongwen Ren, Chao Wang, Peng Du, Chuan Qin 等AAAI 2026
- JAKET: Joint Pre-training of Knowledge Graph and Language UnderstandingDonghan Yu, Chenguang Zhu, Yiming Yang, Michael ZengAAAI 2022 · 被引用 171 次
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu 等ACL 2023 · 被引用 9 次
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng 等ICLR 2022 · 被引用 128 次
