BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, Rui Yan
Abstract
Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery. However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge. To address these issues, we propose BioT5, a comprehensive pre-training framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations. BioT5 utilizes SELFIES for 100% robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature. Furthermore, BioT5 distinguishes between structured and unstructured knowledge, leading to more effective utilization of information. After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities. Our code is available at https://github.com/QizhiPei/BioT5.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Rethinking Text-based Protein Understanding: Retrieval or LLM?Juntong Wu, Zijing Liu, He Cao, Li Hao et al.EMNLP 2025 · 7 citations
- Data-Efficient Molecular Generation with Hierarchical Textual InversionSeojin Kim, Jaehyun Nam, Sihyun Yu, Younghoon Shin et al.ICML 2024 · 6 citations
- DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data VisualizationZhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Jason Zhang et al.ICDE 2025 · 3 citations
- When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMsBogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev et al.ICML 2026 · 3 citations
- Improving Chemical Understanding of LLMs via SMILES ParsingYunhui Jang, Jaehyung Kim, Sungsoo AhnEMNLP 2025 · 2 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby et al.ICLR 2022 · 440 citations
- Motif-based Graph Self-Supervised Learning for Molecular Property PredictionZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.NeurIPS 2021 · 385 citations
Related papers
- 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text ModelingQizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu et al.ICLR 2025
- How to Make Large Language Models Generate 100% Valid Molecules?Wen Tao, Jing Tang, Alvin Chan, Bryan Hooi et al.EMNLP 2025
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke et al.EMNLP 2022 · 112 citations
- MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionJun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park et al.EMNLP 2024 · 2 citations
- Learning Multi-view Molecular Representations with Structured and Unstructured KnowledgeYizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu et al.KDD 2024 · 9 citations
