BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, Rui Yan
摘要
Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery. However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge. To address these issues, we propose BioT5, a comprehensive pre-training framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations. BioT5 utilizes SELFIES for 100% robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature. Furthermore, BioT5 distinguishes between structured and unstructured knowledge, leading to more effective utilization of information. After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities. Our code is available at https://github.com/QizhiPei/BioT5.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Rethinking Text-based Protein Understanding: Retrieval or LLM?Juntong Wu, Zijing Liu, He Cao, Li Hao 等EMNLP 2025 · 被引用 7 次
- Data-Efficient Molecular Generation with Hierarchical Textual InversionSeojin Kim, Jaehyun Nam, Sihyun Yu, Younghoon Shin 等ICML 2024 · 被引用 6 次
- DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data VisualizationZhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Jason Zhang 等ICDE 2025 · 被引用 3 次
- When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMsBogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev 等ICML 2026 · 被引用 3 次
- Improving Chemical Understanding of LLMs via SMILES ParsingYunhui Jang, Jaehyung Kim, Sungsoo AhnEMNLP 2025 · 被引用 2 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby 等ICLR 2022 · 被引用 440 次
- Motif-based Graph Self-Supervised Learning for Molecular Property PredictionZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu 等NeurIPS 2021 · 被引用 385 次
相关 Paper
- 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text ModelingQizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu 等ICLR 2025
- How to Make Large Language Models Generate 100% Valid Molecules?Wen Tao, Jing Tang, Alvin Chan, Bryan Hooi 等EMNLP 2025
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke 等EMNLP 2022 · 被引用 112 次
- MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionJun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park 等EMNLP 2024 · 被引用 2 次
- Learning Multi-view Molecular Representations with Structured and Unstructured KnowledgeYizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu 等KDD 2024 · 被引用 9 次
