Reward Model Routing in Alignment
Xinle Wu, Yao Lu
Abstract
Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single reward model (RM), limiting alignment quality and risking overfitting. Recent work explores RM routing--dynamically selecting an RM from a candidate pool to exploit complementary strengths while maintaining RM calls--but existing methods suffer from cold-start and insufficient exploration. We propose BayesianRouter, a hybrid routing framework that combines offline RM strengths learning with online Bayesian selection. In the offline stage, a multi-task router is trained on preference data to estimate per-RM reliability. In the online stage, a Bayesian Thompson sampling router performs per-query RM selection, initializing RM-specific weight vectors with offline embeddings as Gaussian priors and adaptively updating their posteriors with online rewards to adapt to the evolving policy distribution. Extensive experiments on instruction-following (AlpacaEval-2, Arena-Hard, MT-Bench) and reasoning (GSM8K, MMLU) benchmarks show that BayesianRouter consistently outperforms individual RMs, RM ensembling, and existing routing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- AdaJudge: Adaptive Multi-Perspective Judging for Reward ModelingYongliang Miao, Yangyang Liang, Mengnan DuACL 2026 · 1 citation
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM RoutingYichi Zhang, Fangzheng Xie, Shu Yang, Chong WuICLR 2026 · 1 citation
- From Individual to Common: An Early Exploration of Consensus in Non-verifiable Data for Balanced Preference OptimizationShangjian Yin, Zhouxing ShiACL 2026
Builds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
Related papers
- Ask a Strong LLM Judge when Your Reward Model is UncertainZhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu et al.NeurIPS 2025 · 14 citations
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement LearningHaozhen Zhang, Tao Feng, Jiaxuan YouNeurIPS 2025 · 81 citations
- RLTHF: Targeted Human Feedback for LLM AlignmentYifei Xu, Tusher Chakraborty, Emre Kiciman, Bibek Aryal et al.ICML 2025
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar et al.ACL 2025 · 23 citations
