Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
Wei Wu, Chao Wang, Liyi Chen, Mingze Yin, Yiheng Zhu, Kun Fu, Jieping Ye, Hui Xiong, Zheng Wang
Abstract
Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc49062f-137a-4da8-8f58-e1e177eed071Cited by top-tier papers6
- Graph is a Substrate Across Data ModalitiesZiming Li, Xiao-Ming Wu, Zehong Wang, Jiazheng Li et al.ICML 2026 · 16 citations
- FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM TokensChao Wang, Yixin Song, Jinhui Ye, Chuan Qin et al.NeurIPS 2025 · 7 citations
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang et al.EMNLP 2025 · 2 citations
- Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning ModelsWei Wu, Liyi Chen, Congxi Xiao, Tianfu Wang et al.ACL 2026 · 1 citation
- TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable PromptingJiaming Leng, Yunying Bi, Chuan Qin, Zhenya Huang et al.ACL 2026
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
Related papers
- ProtChatGPT: Towards Understanding Proteins with Hybrid Representation and Large Language ModelsChao Wang, Hehe Fan, Ruijie Quan, Lina Yao et al.SIGIR 2025 · 7 citations
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu et al.ICLR 2024 · 137 citations
- InstructProtein: Aligning Human and Protein Language via Knowledge InstructionZeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin et al.ACL 2024
- SaProt: Protein Language Modeling with Structure-aware VocabularyJin Su, Chenchen Han, Yuyang Zhou, Junjie Shan et al.ICLR 2024 · 285 citations
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng et al.ICLR 2022 · 128 citations
