Adaptive Correlated Monte Carlo for Contextual Categorical Sequence Generation
Xinjie Fan, Yizhe Zhang, Zhendong Wang, Mingyuan Zhou
Abstract
Sequence generation models are commonly refined with reinforcement learning over user-defined metrics. However, high gradient variance hinders the practical use of this method. To stabilize this method for contextual generation of categorical sequences, we estimate the gradient by evaluating a set of correlated Monte Carlo rollouts. Due to the correlation, the number of unique rollouts is random and adaptive to model uncertainty; those rollouts naturally become baselines for each other, and hence are combined to effectively reduce gradient variance. We also demonstrate the use of correlated MC rollouts for binary-tree softmax models which reduce the high generation cost in large vocabulary scenarios, by decomposing each categorical action into a sequence of binary actions. We evaluate our methods on both neural program synthesis and image captioning. The proposed methods yield lower gradient variance and consistent improvement over related baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 969feeee-9141-4594-a9fa-79dbd0fce54dCited by top-tier papers2
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 78 citations
- Recurrent Hierarchical Topic-Guided RNN for Language GenerationDandan Guo, Bo Chen, Ruiying Lu, Mingyuan ZhouICML 2020 · 20 citations
Builds on1
Related papers
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage RankingHaruka Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman et al.ICML 2026
- Reinforcing an Image Caption Generator Using Off-Line Human FeedbackPaul Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han et al.AAAI 2020 · 24 citations
- CARMS: Categorical-Antithetic-REINFORCE Multi-Sample Gradient EstimatorAlek Dimitriev, Mingyuan ZhouNeurIPS 2021 · 10 citations
- The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RLYingru Li, Jiawei Xu, Ziniu Li, Jiacai Liu et al.ICML 2026 · 4 citations
