ReNovo: Retrieval-Based De Novo Mass Spectrometry Peptide Sequencing
Shaorong Chen, Jun Xia, Jingbo Zhou, Lecheng Zhang, Zhangyang Gao, Bozhen Hu, Cheng Tan, Wenjie Du, Stan Z. Li
摘要
Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate de novo peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in de novo peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel Retrieval-based De Novo peptide sequencing methodology, termed ReNovo, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReNovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2021 · 被引用 323 次
相关 Paper
- Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovoJun Xia, Sizhe Liu, Jingbo Zhou, Shaorong Chen 等ICLR 2025
- De novo mass spectrometry peptide sequencing with a transformer modelMelih Yilmaz, William Fondrie, Wout Bittremieux, Sewoong Oh 等ICML 2022 · 被引用 73 次
- Latent Imputation before Prediction: A New Computational Paradigm for De Novo Peptide SequencingYe Du, Chen Yang, Nanxi Yu, Wanyu Lin 等ICML 2025
- Universal Biological Sequence Reranking for Improved De Novo Peptide SequencingZijie Qiu, Jiaqi Wei, Xiang Zhang, Sheng Xu 等ICML 2025
- AdaNovo: Towards Robust De Novo Peptide Sequencing in Proteomics against Data BiasesJun Xia, Shaorong Chen, Jingbo Zhou, Xiaojun Shan 等NeurIPS 2024 · 被引用 5 次
