InstructProtein: Aligning Human and Protein Language via Knowledge Instruction
Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, Huajun Chen
Abstract
Large Language Models (LLMs) have revolutionized the field of natural language processing, but they fall short in comprehending biological sequences such as proteins. To address this challenge, we propose InstructProtein, an innovative LLM that possesses bidirectional generation capabilities in both human and protein languages: (i) taking a protein sequence as input to predict its textual function description and (ii) using natural language to prompt protein sequence generation. To achieve this, we first pre-train an LLM on both protein and natural language corpora, enabling it to comprehend individual languages. Then supervised instruction tuning is employed to facilitate the alignment of these two distinct languages. Herein, we introduce a knowledge graph-based instruction generation framework to construct a high-quality instruction dataset, addressing annotation imbalance and instruction deficits in existing protein-text corpus. In particular, the instructions inherit the structural relations between proteins and function annotations in knowledge graphs, which empowers our model to engage in the causal modeling of protein functions, akin to the chain-of-thought processes in natural languages. Extensive experiments on bidirectional protein-text generation tasks show that InstructProtein outperforms state-of-the-art LLMs by large margins. Moreover, InstructProtein serves as a pioneering step towards text-based protein function prediction and sequence design, effectively bridging the gap between protkbein and human language understanding. 1. We propose InstructProtein, an innovative LLM that enables bidirectional generation between protein and human languages, effectively filling the gap between the two languages. 2. We introduce a protein instruction generation framework with knowledge graphs, resulting in the first high-quality protein instruction dataset for tuning LLMs. 3. The InstructProtein outperforms state-of-the-art LLMs by a substantial margin, serving as a pioneering step toward text-guided protein function prediction and sequence design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f50dca23-1994-4943-b8f0-eee92a64a3aeCited by top-tier papers5
- MMSite: A Multi-modal Framework for the Identification of Active Sites in ProteinsSong Ouyang, Huiyu Cai, Yong Luo, Kehua Su et al.NeurIPS 2024 · 10 citations
- Controllable Protein Sequence Generation with LLM Preference OptimizationXiangyu Liu, Yi Liu, Silei Chen, Wei HuAAAI 2025 · 8 citations
- Enhancing Safe and Controllable Protein Generation via Knowledge Preference OptimizationYuhao Wang, Keyan Ding, Kehua Feng, Zeyuan Wang et al.ACL 2025 · 2 citations
- Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMsWei Wu, Chao Wang, Liyi Chen, Mingze Yin et al.KDD 2025 · 1 citation
- Logical Consistency of Large Language Models in Fact-CheckingBishwamittra Ghosh, Sarah Hasan, Naheed Anjum Arafat, Arijit KhanICLR 2025
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu et al.ICLR 2024 · 137 citations
- Multi-level Protein Structure Pre-training via Prompt LearningZeyuan Wang, Qiang Zhang, Shuangwei Hu, Haoran Yu et al.ICLR 2023
- Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive AlignmentXiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine et al.NeurIPS 2025 · 12 citations
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng et al.ICLR 2022 · 128 citations
- ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language ModelsChao Wang, Hehe Fan, Ruijie Quan, Lina Yao et al.SIGIR 2025 · 7 citations
