Learning to Answer from Correct Demonstrations
Nirmit Joshi, Gene Li, Siddharth Bhandari, Shiva Prasad Kasiviswanathan, Cong Ma, Nathan Srebro
摘要
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at test time. Learning is based on demonstrations of some correct answer to each training question, as in Supervised Fine Tuning (SFT). We formalize the problem as imitation learning (i.e., apprenticeship learning) in contextual bandits, with offline demonstrations from some expert (optimal, or very good) policy, without explicitly observed rewards. In contrast to prior work, which assumes the demonstrator belongs to a bounded-complexity policy class, we propose relying only on the underlying reward model (i.e., specifying which answers are correct) being in a bounded-complexity class, which we argue is a strictly weaker assumption. We show that likelihood-maximization methods can fail in this setting, and instead present an approach that learns to answer nearly as well as the demonstrator, with sample complexity logarithmic in the cardinality of the reward class. Our method is similar to Syed and Schapire 2007, when adapted to a contextual bandit (i.e., single step) setup, but is a simple one-pass online approach that enjoys an ``optimistic rate'' (i.e., when the demonstrator is optimal, versus in Syed and Schapire 2007, and works even with arbitrarily adaptive demonstrations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 被引用 137 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
相关 Paper
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie 等ICLR 2026 · 被引用 3 次
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
- Contextual Bandits and Imitation Learning with Preference-Based Active QueriesAyush Sekhari, Karthik Sridharan, Wen Sun, Runzhe WuNeurIPS 2023 · 被引用 18 次
- Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language ModelsFeng Zhao, Hong Zhang, Yu Yang, Ruilin Zhao 等ICML 2026
- Subject-driven Text-to-Image Generation via Apprenticeship LearningWenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz 等NeurIPS 2023 · 被引用 265 次
