Adaptive Correlated Monte Carlo for Contextual Categorical Sequence Generation
Xinjie Fan, Yizhe Zhang, Zhendong Wang, Mingyuan Zhou
摘要
Sequence generation models are commonly refined with reinforcement learning over user-defined metrics. However, high gradient variance hinders the practical use of this method. To stabilize this method for contextual generation of categorical sequences, we estimate the gradient by evaluating a set of correlated Monte Carlo rollouts. Due to the correlation, the number of unique rollouts is random and adaptive to model uncertainty; those rollouts naturally become baselines for each other, and hence are combined to effectively reduce gradient variance. We also demonstrate the use of correlated MC rollouts for binary-tree softmax models which reduce the high generation cost in large vocabulary scenarios, by decomposing each categorical action into a sequence of binary actions. We evaluate our methods on both neural program synthesis and image captioning. The proposed methods yield lower gradient variance and consistent improvement over related baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 被引用 78 次
- Recurrent Hierarchical Topic-Guided RNN for Language GenerationDandan Guo, Bo Chen, Ruiying Lu, Mingyuan ZhouICML 2020 · 被引用 20 次
它引用的顶会 Paper1
相关 Paper
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage RankingHaruka Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman 等ICML 2026
- Reinforcing an Image Caption Generator Using Off-Line Human FeedbackPaul Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han 等AAAI 2020 · 被引用 24 次
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient EstimatorAlek Dimitriev, Mingyuan ZhouNeurIPS 2021 · 被引用 10 次
- The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RLYingru Li, Jiawei Xu, Ziniu Li, Jiacai Liu 等ICML 2026 · 被引用 4 次
