InstructProtein: Aligning Human and Protein Language via Knowledge Instruction
Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, Huajun Chen
摘要
Large Language Models (LLMs) have revolutionized the field of natural language processing, but they fall short in comprehending biological sequences such as proteins. To address this challenge, we propose InstructProtein, an innovative LLM that possesses bidirectional generation capabilities in both human and protein languages: (i) taking a protein sequence as input to predict its textual function description and (ii) using natural language to prompt protein sequence generation. To achieve this, we first pre-train an LLM on both protein and natural language corpora, enabling it to comprehend individual languages. Then supervised instruction tuning is employed to facilitate the alignment of these two distinct languages. Herein, we introduce a knowledge graph-based instruction generation framework to construct a high-quality instruction dataset, addressing annotation imbalance and instruction deficits in existing protein-text corpus. In particular, the instructions inherit the structural relations between proteins and function annotations in knowledge graphs, which empowers our model to engage in the causal modeling of protein functions, akin to the chain-of-thought processes in natural languages. Extensive experiments on bidirectional protein-text generation tasks show that InstructProtein outperforms state-of-the-art LLMs by large margins. Moreover, InstructProtein serves as a pioneering step towards text-based protein function prediction and sequence design, effectively bridging the gap between protkbein and human language understanding. 1. We propose InstructProtein, an innovative LLM that enables bidirectional generation between protein and human languages, effectively filling the gap between the two languages. 2. We introduce a protein instruction generation framework with knowledge graphs, resulting in the first high-quality protein instruction dataset for tuning LLMs. 3. The InstructProtein outperforms state-of-the-art LLMs by a substantial margin, serving as a pioneering step toward text-guided protein function prediction and sequence design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MMSite: A Multi-modal Framework for the Identification of Active Sites in ProteinsSong Ouyang, Huiyu Cai, Yong Luo, Kehua Su 等NeurIPS 2024 · 被引用 10 次
- Controllable Protein Sequence Generation with LLM Preference OptimizationXiangyu Liu, Yi Liu, Silei Chen, Wei HuAAAI 2025 · 被引用 8 次
- Enhancing Safe and Controllable Protein Generation via Knowledge Preference OptimizationYuhao Wang, Keyan Ding, Kehua Feng, Zeyuan Wang 等ACL 2025 · 被引用 2 次
- Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMsWei Wu, Chao Wang, Liyi Chen, Mingze Yin 等KDD 2025 · 被引用 1 次
- Logical Consistency of Large Language Models in Fact-CheckingBishwamittra Ghosh, Sarah Hasan, Naheed Anjum Arafat, Arijit KhanICLR 2025
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu 等ICLR 2024 · 被引用 137 次
- Multi-level Protein Structure Pre-training via Prompt LearningZeyuan Wang, Qiang Zhang, Shuangwei Hu, Haoran Yu 等ICLR 2023
- Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive AlignmentXiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine 等NeurIPS 2025 · 被引用 12 次
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng 等ICLR 2022 · 被引用 128 次
- ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language ModelsChao Wang, Hehe Fan, Ruijie Quan, Lina Yao 等SIGIR 2025 · 被引用 7 次
