LASeR: Learning to Adaptively Select Reward Models with Multi-Arm Bandits
Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
Abstract
Reward Models (RMs) are crucial to aligning large language models (LLMs), but the degree to which an RM specialized to one task (e.g. writing) generalizes to new tasks (e.g. math) is often not known a priori, often making using only one fixed RM to train LLMs suboptimal. However, optimizing LLMs with multiple RMs simultaneously can incur a prohibitively high computational cost and lead to conflicting signals from different RMs that may degrade performance. To address these challenges, we introduce LASER (Learning to Adaptively Select Rewards), which frames reward model selection as a multi-armed bandit problem, efficiently and iteratively training LLMs using multiple RMs by selecting the most wellsuited RM for each instance. On commonsense and math reasoning tasks, we show that LASER boosts iterative LLM training, improving the absolute average accuracy of Llama-3-8B over three datasets by 2.67% over an ensemble of RM scores while also showing superior efficiency (e.g., a 2× speedup). Moreover, on WildChat (open-ended instruction-following tasks), LASER leads to a 72.69% AlpacaEval win rate over the RM score ensemble baseline. Extending to longcontext generation, LASER improves by 2.96 F1 points (avg.) on single-document QA tasks and 2.97 F1 points on few-shot learning over the RM score ensemble baseline with best-of-n sampling. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on40
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Mutual-Taught for Co-adapting Policy and Reward ModelsTianyuan Shi, Canbin Huang, Fanqi Wan, Longguang Zhong et al.ACL 2025 · 1 citation
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin et al.ICLR 2026 · 147 citations
- Language Imbalance Driven Rewarding for Multilingual Self-improvingWen Yang, Junhong Wu, Chen Wang, Chengqing Zong et al.ICLR 2025
- A Systematic Analysis of Base Model Choice for Reward ModelingKian Ahrabian, Pegah Jandaghi, Negar Mokhberian, Sai Praneeth Karimireddy et al.EMNLP 2025
- AgentRM: Enhancing Agent Generalization with Reward ModelingYu Xia, Jingru Fan, Weize Chen, Siyu Yan et al.ACL 2025 · 20 citations
