Learning Multi-view Molecular Representations with Structured and Unstructured Knowledge
Yizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu, Zikun Nie, Hao Zhou, Zaiqing Nie
Abstract
Capturing molecular knowledge with representation learning approaches holds significant potential in vast scientific fields such as chemistry and life science. An effective and generalizable molecular representation is expected to capture the consensus and complementary molecular expertise from diverse views and perspectives. However, existing works fall short in learning multi-view molecular representations, due to challenges in explicitly incorporating view information and handling molecular knowledge from heterogeneous sources. To address these issues, we present MV-Mol, a molecular representation learning model that harvests multi-view molecular expertise from chemical structures, unstructured knowledge from biomedical texts, and structured knowledge from knowledge graphs. We utilize text prompts to model view information and design a fusion architecture to extract view-based molecular representations. We develop a two-stage pre-training procedure, exploiting heterogeneous data of varying quality and quantity. Through extensive experiments, we show that MV-Mol provides improved representations that substantially benefit molecular property prediction. Additionally, MV-Mol exhibits state-of-the-art performance in multi-modal comprehension of molecular structures and texts. Code and data are available at https://github.com/PharMolix/OpenBioMed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca5ff250-481f-433b-b94f-5bca7e2dd037Cited by top-tier papers3
- MedAlign: Enhancing Combinatorial Medication Recommendation with Multi-modality AlignmentHang Lv, Zixuan Guo, Zijie Wu, Yanchao Tan et al.ACM MM 2025 · 4 citations
- TopoFormer: Topology Meets Attention for Graph LearningMd Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora, Baris CoskunuzerICLR 2026 · 2 citations
- Advancing Molecular Graph-Text Pre-training via Fine-grained AlignmentYibo Li, Yuan Fang, Mengmei Zhang, Chuan ShiKDD 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- Cross-view Contrastive Unification Guides Generative Pretraining for Molecular Property PredictionJunyu Lin, Yan Zheng, Xinyue Chen, Yazhou Ren et al.ACM MM 2024 · 8 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
- Unified 2D and 3D Pre-Training of Molecular RepresentationsJinhua Zhu, Yingce Xia, Lijun Wu, Shufang Xie et al.KDD 2022 · 53 citations
- Bi-level Contrastive Learning for Knowledge-Enhanced Molecule RepresentationsPengcheng Jiang, Cao Xiao, Tianfan Fu, Parminder Bhatia et al.AAAI 2025 · 7 citations
- DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality ExpertsMingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu et al.ACM MM 2025
