Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing
Xiang Zhang, Jiaqi Wei, Zijie Qiu, Sheng Xu, Nanqing Dong, Zhiqiang Gao, Siqi Sun
Abstract
Peptide sequencing-the process of identifying amino acid sequences from mass spectrometry data-is a fundamental task in proteomics. Non-Autoregressive Transformers (NATs) have proven highly effective for this task, outperforming traditional methods. Unlike autoregressive models, which generate tokens sequentially, NATs predict all positions simultaneously, leveraging bidirectional context through unmasked self-attention. However, existing NAT approaches often rely on Connectionist Temporal Classification (CTC) loss, which presents significant optimization challenges due to CTC's complexity and increases the risk of training failures. To address these issues, we propose an improved non-autoregressive peptide sequencing model that incorporates a structured protein sequence curriculum learning strategy. This approach adjusts protein's learning difficulty based on the model's estimated protein generational capabilities through a sampling process, progressively learning peptide generation from simple to complex sequences. Additionally, we introduce a self-refining inference-time module that iteratively enhances predictions using learned NAT token embeddings, improving sequence accuracy at a fine-grained level. Our curriculum learning strategy reduces NAT training failures frequency by more than 90% based on sampled training over various data distributions. Evaluations on nine benchmark species demonstrate that our approach outperforms all previous methods across multiple metrics and species. Model and source code are available at Github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88966d82-3f53-41f2-84a0-b8437a95406dCited by top-tier papers1
Ask how each one uses itBuilds on19
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Step-unrolled Denoising Autoencoders for Text GenerationNikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen et al.ICLR 2022 · 142 citations
- Structure-informed Language Models Are Protein DesignersZaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou et al.ICML 2023 · 130 citations
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
Related papers
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin et al.AAAI 2020 · 91 citations
- De novo mass spectrometry peptide sequencing with a transformer modelMelih Yilmaz, William Fondrie, Wout Bittremieux, Sewoong Oh et al.ICML 2022 · 73 citations
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun et al.ACL 2023 · 4 citations
- Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass ControlShaorong Chen, Jingbo Zhou, Jun XiaAAAI 2026
- ContraNovo: A Contrastive Learning Approach to Enhance De Novo Peptide SequencingZhi Jin, Sheng Xu, Xiang Zhang, Tianze Ling et al.AAAI 2024 · 31 citations
