3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text Modeling
Qizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu, Lijun Wu
摘要
The integration of molecular and natural language representations has emerged as a focal point in molecular science, with recent advancements in Language Models (LMs) demonstrating significant potential for comprehensive modeling of both domains. However, existing approaches face notable limitations, particularly in their neglect of three-dimensional (3D) information, which is crucial for understanding molecular structures and functions. While some efforts have been made to incorporate 3D molecular information into LMs using external structure encoding modules, significant difficulties remain, such as insufficient interaction across modalities in pre-training and challenges in modality alignment. To address the limitations, we propose 3D-MolT5, a unified framework designed to model molecule in both sequence and 3D structure spaces. The key innovation of our approach lies in mapping fine-grained 3D substructure representations into a specialized 3D token vocabulary. This methodology facilitates the seamless integration of sequence and structure representations in a tokenized format, enabling 3D-MolT5 to encode molecular sequences, molecular structures, and text sequences within a unified architecture. Leveraging this tokenized input strategy, we build a foundation model that unifies the sequence and structure data formats. We then conduct joint pretraining with multi-task objectives to enhance the model's comprehension of these diverse modalities within a shared representation space. Thus, our approach significantly improves cross-modal interaction and alignment, addressing key challenges in previous work. Further instruction tuning demonstrated that our 3D-MolT5 has strong generalization ability and surpasses existing methods with superior performance in multiple downstream tasks, such as nearly 70% improvement on the molecular property prediction task compared to state-of-the-art methods. Our code is available at https://github.com/QizhiPei/3D-MolT5 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular UnderstandingZihao Jing, QIUHAO Zeng, Ruiyi Fang, Yan Sun 等ICLR 2026 · 被引用 3 次
- RTMol: Rethinking Molecule-text Alignment in a Round-trip ViewLetian Chen, Runhan Shi, Gufeng Yu, Yang YangAAAI 2026 · 被引用 1 次
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng 等ICLR 2023 · 被引用 254 次
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu 等ICLR 2024 · 被引用 137 次
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke 等EMNLP 2022 · 被引用 112 次
相关 Paper
- Towards 3D Molecule-Text Interpretation in Language ModelsSihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang 等ICLR 2024 · 被引用 87 次
- DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality ExpertsMingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 等ACM MM 2025
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsQizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu 等EMNLP 2023 · 被引用 40 次
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
- Advancing Molecular Graph-Text Pre-training via Fine-grained AlignmentYibo Li, Yuan Fang, Mengmei Zhang, Chuan ShiKDD 2025
