Learning to Rank Generation with Pairwise Partial Rewards
Youngwon Lee, Jinu Lee, Seung-won Hwang
摘要
This paper studies the use of reinforcement learning for conditional text generation, which overcomes the limitation of the prevalent supervised maximum likelihood estimation approach. However, it still suffers from challenges including the large action space and the delayed reward, as the reward can be computed only after an entire sequence is generated. To address these challenges, we propose a method that provides partial rewards for intermediate actions taken on partial sequences. This enables the model to promptly prioritize actions that lead to the generation of more desirable sequences. Our method's key contribution lies in its focus on distinguishing relatively more desirable actions rather than striving to precisely estimate pointwise values for arbitrary partial sequences. Instead, our reward shaping method learns to discern the relative desirability between pairs of actions, or rank actions in a pairwise manner, only when necessary and feasible. This is materialized in an efficient way by leveraging the prefix tree constructed from the sampled sequences. Experimental results on paraphrase generation and constrained machine translation tasks showcase the effectiveness of our method. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 被引用 329 次
相关 Paper
- Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text GenerationWenqing Chen, Jidong Tian, Caoyun Fan, Yitian Li 等AAAI 2023 · 被引用 2 次
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 被引用 24 次
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 被引用 69 次
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 被引用 7 次
- Goal-directed Generation of Discrete Structures with Conditional Generative ModelsAmina Mollaysa, Brooks Paige, Alexandros KalousisNeurIPS 2020 · 被引用 12 次
