Retrieval-Augmented Language Model for Knowledge-aware Protein Encoding
Jiasheng Zhang, Delvin Ce Zhang, Shuang Liang, Zhengpin Li, Rex Ying, Jie Shao
Abstract
Protein language models often struggle to capture biological functions due to their lack of factual knowledge (e.g., gene descriptions). Existing solutions leverage protein knowledge graphs (PKGs) as auxiliary pre-training objectives, but lack explicit integration of task-oriented knowledge, making them suffer from limited knowledge exploitation and catastrophic forgetting. The root cause is that they fail to align PKGs with task-specific data, forcing their knowledge modeling to adapt to the knowledge-isolated nature of downstream tasks. In this paper, we propose Knowledge-aware retrieval augmented protein language model (Kara), achieving the first task-oriented and explicit integration of PKGs and protein language models. With a knowledge retriever learning to predict linkages between PKG and task proteins, Kara unifies the knowledge integration of the pre-training and fine-tuning stages with a structure-based regularization, mitigating catastrophic forgetting. To ensure task-oriented integration, Kara uses contextualized virtual tokens to extract graph context as task-specific knowledge for new proteins. Experiments show that Kara outperforms existing knowledge-enhanced models in 6 representative tasks, achieving on average 5.1% improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f330e9af-24bd-49eb-bf2b-758c7d82d5a2Builds on9
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question AnsweringXiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla et al.NeurIPS 2024 · 384 citations
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Making Large Language Models Perform Better in Knowledge Graph CompletionYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.ACM MM 2024 · 86 citations
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu et al.AAAI 2022 · 36 citations
Related papers
- Protein Representation Learning via Knowledge Enhanced Primary Structure ReasoningHong-Yu Zhou, Yunxiang Fu, Zhicheng Zhang, Cheng Bian et al.ICLR 2023
- Enhancing Conversational Recommender Systems with Tree-Structured Knowledge and Pretrained Language ModelsYongwen Ren, Chao Wang, Peng Du, Chuan Qin et al.AAAI 2026
- JAKET: Joint Pre-training of Knowledge Graph and Language UnderstandingDonghan Yu, Chenguang Zhu, Yiming Yang, Michael ZengAAAI 2022 · 171 citations
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu et al.ACL 2023 · 9 citations
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng et al.ICLR 2022 · 128 citations
