ReNovo: Retrieval-Based De Novo Mass Spectrometry Peptide Sequencing
Shaorong Chen, Jun Xia, Jingbo Zhou, Lecheng Zhang, Zhangyang Gao, Bozhen Hu, Cheng Tan, Wenjie Du, Stan Z. Li
Abstract
Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate de novo peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in de novo peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel Retrieval-based De Novo peptide sequencing methodology, termed ReNovo, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReNovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2021 · 323 citations
Related papers
- Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovoJun Xia, Sizhe Liu, Jingbo Zhou, Shaorong Chen et al.ICLR 2025
- De novo mass spectrometry peptide sequencing with a transformer modelMelih Yilmaz, William Fondrie, Wout Bittremieux, Sewoong Oh et al.ICML 2022 · 73 citations
- Latent Imputation before Prediction: A New Computational Paradigm for De Novo Peptide SequencingYe Du, Chen Yang, Nanxi Yu, Wanyu Lin et al.ICML 2025
- Universal Biological Sequence Reranking for Improved De Novo Peptide SequencingZijie Qiu, Jiaqi Wei, Xiang Zhang, Sheng Xu et al.ICML 2025
- AdaNovo: Towards Robust De Novo Peptide Sequencing in Proteomics against Data BiasesJun Xia, Shaorong Chen, Jingbo Zhou, Xiaojun Shan et al.NeurIPS 2024 · 5 citations
