SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
Chris Cundy, Stefano Ermon
摘要
In many domains, autoregressive models can attain high likelihood on the task of predicting the next observation. However, this maximum-likelihood (MLE) objective does not necessarily match a downstream use-case of autoregressively generating high-quality sequences. The MLE objective weights sequences proportionally to their frequency under the data distribution, with no guidance for the model's behaviour out of distribution (OOD): leading to compounding error during autoregressive generation. In order to address this compounding error problem, we formulate sequence generation as an imitation learning (IL) problem. This allows us to minimize a variety of divergences between the distribution of sequences generated by an autoregressive model and sequences from a dataset, including divergences with weight on OOD generated sequences. The IL framework also allows us to incorporate backtracking by introducing a backspace action into the generation process. This further mitigates the compounding error problem by allowing the model to revert a sampled token if it takes the sequence OOD. Our resulting method, SequenceMatch, can be implemented without adversarial training or architectural changes. We identify the SequenceMatch- divergence as a more suitable training objective for autoregressive models which are used for generation. We show that empirically, SequenceMatch training leads to improvements over MLE on text generation with language models and arithmetic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-TuningGokul Swamy, Sanjiban Choudhury, Wen Sun, Steven Wu 等ICLR 2026 · 被引用 66 次
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making AbilitiesThomas Schmied, Jörg Bornschein, Jordi Grau-Moya, Markus Wulfmeier 等ICLR 2026 · 被引用 41 次
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja 等NeurIPS 2024 · 被引用 26 次
- Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsBo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao 等NeurIPS 2025 · 被引用 23 次
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper7
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- A Theoretical Analysis of the Repetition Problem in Text GenerationZihao Fu, Wai Lam, Anthony Man-Cho So, Bei ShiAAAI 2021 · 被引用 114 次
相关 Paper
- Time-series Generation by Contrastive ImitationDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2021 · 被引用 29 次
- DistiLLM: Towards Streamlined Distillation for Large Language ModelsJongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young YunICML 2024 · 被引用 86 次
- An Imitation Learning Curriculum for Text Editing with Non-Autoregressive ModelsSweta Agrawal, Marine CarpuatACL 2022
- MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingSean Welleck, Kyunghyun ChoAAAI 2021 · 被引用 8 次
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu 等EMNLP 2020 · 被引用 11 次
