Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries
Carl Edwards, ChengXiang Zhai, Heng Ji
Abstract
We propose a new task, Text2Mol, to retrieve molecules using natural language descriptions as queries. Natural language and molecules encode information in very different ways, which leads to the exciting but challenging problem of integrating these two very different modalities. Although some work has been done on text-based retrieval and structurebased retrieval, this new task requires integrating molecules and natural language more directly. Moreover, this can be viewed as an especially challenging cross-lingual retrieval problem by considering the molecules as a language with a very unique grammar. We construct a paired dataset of molecules and their corresponding text descriptions, which we use to learn an aligned common semantic embedding space for retrieval. We extend this to create a cross-modal attention-based model for explainability and reranking by interpreting the attentions as association rules. We also employ an ensemble approach to integrate our different architectures, which significantly improves results from 0.372 to 0.499 MRR. This new multimodal approach opens a new perspective on solving problems in chemistry literature understanding and molecular machine learning. 1 1 The programs and data are publicly available at github.com/cnedwards/text2mol for research purposes. Water is an oxygen hydride consisting of an oxygen atom that is covalently bonded to two hydrogen atoms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d09f5386-a0b3-4e7f-b0b6-600aff5f82dcCited by top-tier papers45
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Unifying Molecular and Textual Representations via Multi-task Language ModellingDimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther et al.ICML 2023 · 126 citations
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke et al.EMNLP 2022 · 112 citations
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
- Towards 3D Molecule-Text Interpretation in Language ModelsSihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang et al.ICLR 2024 · 87 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead et al.ACL 2020 · 87 citations
- Joint Biomedical Entity and Relation Extraction with Knowledge-Enhanced Collective InferenceTuan Manh Lai, Heng Ji, ChengXiang Zhai, Quan Hung TranACL 2021
- Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning RepresentationZixuan Zhang, Nikolaus Nova Parulian, Heng Ji, Ahmed Elsayed et al.ACL 2021
Related papers
- Predictive Chemistry Augmented with Text RetrievalYujie Qian, Zhening Li, Zhengkai Tu, Connor W. Coley et al.EMNLP 2023 · 11 citations
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaJunxian Li, Di Zhang, Xunzhi Wang, Zeying Hao et al.AAAI 2025 · 71 citations
- Improving Large Molecular Language Model via Relation-aware Multimodal CollaborationJinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. KimAAAI 2026
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song et al.NeurIPS 2025 · 5 citations
- UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image CandidatesYupei Yang, Lin Yang, Wanxi Deng, Lin Qu et al.KDD 2026
