Translation between Molecules and Natural Language
Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, Heng Ji
摘要
We present MolT5 -a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings. MolT5 allows for new, useful, and challenging analogs of traditional vision-language tasks, such as molecule captioning and text-based de novo molecule generation (altogether: translation between molecules and language), which we explore for the first time. Since MolT5 pretrains models on single-modal data, it helps overcome the chemistry domain shortcoming of data scarcity. Furthermore, we consider several metrics, including a new cross-modal embedding-based metric, to evaluate the tasks of molecule captioning and text-based molecule generation. Our results show that MolT5-based models are able to generate outputs, both molecules and captions, which in many cases are high quality 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- Grammar Prompting for Domain-Specific Language Generation with Large Language ModelsBailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao 等NeurIPS 2023 · 被引用 138 次
- Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language ModelsYin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu 等ICLR 2024 · 被引用 137 次
- Unifying Molecular and Textual Representations via Multi-task Language ModellingDimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther 等ICML 2023 · 被引用 126 次
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu 等NeurIPS 2023 · 被引用 97 次
- Towards 3D Molecule-Text Interpretation in Language ModelsSihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang 等ICLR 2024 · 被引用 87 次
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Scaling Up Vision-Language Pretraining for Image CaptioningXiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang 等CVPR 2022 · 被引用 203 次
- Chemical-Reaction-Aware Molecule Representation LearningHongwei Wang, Weijiang Li, Xiaomeng Jin, Kyunghyun Cho 等ICLR 2022 · 被引用 79 次
相关 Paper
- 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text ModelingQizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu 等ICLR 2025
- MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionJun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park 等EMNLP 2024 · 被引用 2 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsQizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu 等EMNLP 2023 · 被引用 40 次
- RTMol: Rethinking Molecule-text Alignment in a Round-trip ViewLetian Chen, Runhan Shi, Gufeng Yu, Yang YangAAAI 2026 · 被引用 1 次
