Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra
Yiwen Zhang, Keyan Ding, Yihang Wu, Xiang Zhuang, Yi Yang, Qiang Zhang, Huajun Chen
摘要
Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- DiffMS: Diffusion Generation of Molecules Conditioned on Mass SpectraMontgomery Bohde, Mrunali Manjrekar, Runzhong Wang, Shuiwang Ji 等ICML 2025
- MADGEN: Mass-Spec attends to De Novo Molecular generationYinkai Wang, Xiaohui Chen, Liping Liu, Soha HassounICLR 2025
相关 Paper
- MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure ElucidationYang Han, Pengyu Wang, Kai Yu, Xin Chen 等NeurIPS 2025 · 被引用 11 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
- FreeRet: MLLMs as Training-Free RetrieversYuhan Zhu, Xiangyu Zeng, Chenting Wang, Xinhao Li 等ICML 2026 · 被引用 5 次
- Dual-Branch Multi-Granularity Network with Structured Contrastive Ranking for Cross-Modal RetrievalZihao Chen, Chenyang Bu, Shengwei Ji, Xindong WuWWW 2026
- FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference TimeMontgomery Bohde, Hongxuan Liu, Mrunali Manjrekar, Magdalena Lederbauer 等ICML 2026 · 被引用 3 次
