DORB: Dynamically Optimizing Multiple Rewards with Bandits
Ramakanth Pasunuru, Han Guo, Mohit Bansal
摘要
Policy gradients-based reinforcement learning has proven to be a promising approach for directly optimizing non-differentiable evaluation metrics for language generation tasks. However, optimizing for a specific metric reward leads to improvements in mostly that metric only, suggesting that the model is gaming the formulation of that metric in a particular way without often achieving real qualitative improvements. Hence, it is more beneficial to make the model optimize multiple diverse metric rewards jointly. While appealing, this is challenging because one needs to manually decide the importance and scaling weights of these metric rewards. Further, it is important to consider using a dynamic combination and curriculum of metric rewards that flexibly changes over time. Considering the above aspects, in our work, we automate the optimization of multiple metric rewards simultaneously via a multi-armed bandit approach (DORB), where at each round, the bandit chooses which metric reward to optimize next, based on expected arm gains. We use the Exp3 algorithm for bandits and formulate two approaches for bandit rewards: (1) Single Multi-reward Bandit (SM-Bandit); (2) Hierarchical Multi-reward Bandit (HM-Bandit). We empirically show the effectiveness of our approaches via various automatic metrics and human evaluation on two important NLG tasks: question generation and data-to-text generation. Finally, we present interpretable analyses of the learned bandit curriculum over the optimized rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu 等ICLR 2024 · 被引用 124 次
- REAL: Regression-Aware Reinforcement Learning for LLM-as-a-JudgeYasi Zhang, Tianyu Chen, Mingyuan Zhou, Oscar Leong 等ICML 2026
- Dynamic Multi-Reward Weighting for Multi-Style Controllable GenerationKarin de Langis, Ryan Koo, Dongyeop KangEMNLP 2024 · 被引用 3 次
- Efficient Multi-objective Prompt Optimization via Pure-exploration BanditsDonghao Li, Chengshuai Shi, Weijuan Ou, Cong Shen 等ICLR 2026 · 被引用 2 次
