Multilingual Molecular Representation Learning via Contrastive Pre-training
Zhihui Guo, Pramod Kumar Sharma, Andy Martinez, Liang Du, Robin Abraham
Abstract
Molecular representation learning plays an essential role in cheminformatics. Recently, language model-based approaches have gained popularity as an alternative to traditional expert-designed features to encode molecules. However, these approaches only utilize a single molecular language for representation learning. Motivated by the fact that a given molecule can be described using different languages such as Simplified Molecular Line Entry System (SMILES), The International Union of Pure and Applied Chemistry (IUPAC), and The IUPAC International Chemical Identifier (InChI), we propose a multilingual molecular embedding generation approach called MM-Deacon (multilingual molecular domain embedding analysis via contrastive learning). MM-Deacon is pre-trained using SMILES and IUPAC as two different languages on large-scale molecules. We evaluated the robustness of our method on seven molecular property prediction tasks from MoleculeNet benchmark, zero-shot cross-lingual retrieval, and a drug-drug interaction prediction task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 384871c3-eeae-4bf0-94a3-b431ff064757Cited by top-tier papers5
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguagePhilipp Seidl, Andreu Vall, Sepp Hochreiter, Günter KlambauerICML 2023 · 69 citations
- Fractional Denoising for 3D Molecular Pre-trainingShikun Feng, Yuyan Ni, Yanyan Lan, Zhi-Ming Ma et al.ICML 2023 · 42 citations
- Mol-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test ObjectiveJunwei Yang, Kangjie Zheng, Siyu Long, Zaiqing Nie et al.ICML 2024 · 19 citations
- ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingKangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang et al.ICML 2024 · 17 citations
- MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text PromptsHaoqiang Guo, Sendong Zhao, Haochun Wang, Yanrui Du et al.AAAI 2024 · 17 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby et al.ICLR 2022 · 440 citations
Related papers
- Molecular String Representation Preferences in Pretrained LLMs: A Comparative Study in Zero- & Few-Shot Molecular Property PredictionGeorge Arthur Baker, Mario Sanz-Guerrero, Katharina von der WenseEMNLP 2025
- Chemical-Reaction-Aware Molecule Representation LearningHongwei Wang, Weijiang Li, Xiaomeng Jin, Kyunghyun Cho et al.ICLR 2022 · 79 citations
- MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionJun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park et al.EMNLP 2024 · 2 citations
- MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and GenerationFeiyang Cai, Jiahui Bai, Tao Tang, Guijuan He et al.ICLR 2026 · 10 citations
- Domain-Agnostic Molecular Generation with Chemical FeedbackYin Fang, Ningyu Zhang, Zhuo Chen, Lingbing Guo et al.ICLR 2024 · 33 citations
