ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
Le Zhuo, Zewen Chi, Minghao Xu, Heyan Huang, Jianan Zhao, Heqi Zheng, Conghui He, Xian-Ling Mao, Wentao Zhang
Abstract
We propose PROTLLM, a versatile crossmodal large language model (LLM) for both protein-centric and protein-language tasks. PROTLLM features a unique dynamic protein mounting mechanism, enabling it to handle complex inputs where the natural language text is interspersed with an arbitrary number of proteins. Besides, we propose the proteinas-word language modeling approach to train PROTLLM. By developing a specialized protein vocabulary, we equip the model with the capability to predict not just natural language but also proteins from a vast pool of candidates. Additionally, we construct a largescale interleaved protein-text dataset, named InterPT, for pre-training. This dataset comprehensively encompasses both (1) structured data sources like protein annotations and (2) unstructured data sources like biological research papers, thereby endowing PROTLLM with crucial knowledge for understanding proteins. We evaluate PROTLLM on classic supervised protein-centric tasks and explore its novel protein-language applications. Experimental results demonstrate that PROTLLM not only achieves superior performance against proteinspecialized baselines on protein-centric tasks but also induces zero-shot and in-context learning capabilities on protein-language tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c481d2d-5a25-4ac6-aa1f-92c56faea83bCited by top-tier papers7
- Controllable Protein Sequence Generation with LLM Preference OptimizationXiangyu Liu, Yi Liu, Silei Chen, Wei HuAAAI 2025 · 8 citations
- Rethinking Text-based Protein Understanding: Retrieval or LLM?Juntong Wu, Zijing Liu, He Cao, Li Hao et al.EMNLP 2025 · 7 citations
- MutaPLM: Protein Language Modeling for Mutation Explanation and EngineeringYizhen Luo, Zikun Nie, Massimo Hong, Suyuan Zhao et al.NeurIPS 2024 · 6 citations
- Large Language and Protein Assistant for Protein-Protein Interactions PredictionPeng Zhou, Pengsen Ma, Jianmin Wang, Xibao Cai et al.ACL 2025 · 2 citations
- Curriculum Model Merging: Harmonizing Chemical LLMs for Enhanced Cross-Task GeneralizationBaoyi He, Luotian Yuan, Ying Wei, Fei WuNeurIPS 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- ProtCLIP: Function-Informed Protein Multi-Modal LearningHanjing Zhou, Mingze Yin, Wei Wu, Mingyang Li et al.AAAI 2025 · 11 citations
- ProtT3: Protein-to-Text Generation for Text-based Protein UnderstandingZhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang et al.ACL 2024 · 6 citations
- InstructProtein: Aligning Human and Protein Language via Knowledge InstructionZeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin et al.ACL 2024
