Reward Model Routing in Alignment
Xinle Wu, Yao Lu
摘要
Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single reward model (RM), limiting alignment quality and risking overfitting. Recent work explores RM routing--dynamically selecting an RM from a candidate pool to exploit complementary strengths while maintaining RM calls--but existing methods suffer from cold-start and insufficient exploration. We propose BayesianRouter, a hybrid routing framework that combines offline RM strengths learning with online Bayesian selection. In the offline stage, a multi-task router is trained on preference data to estimate per-RM reliability. In the online stage, a Bayesian Thompson sampling router performs per-query RM selection, initializing RM-specific weight vectors with offline embeddings as Gaussian priors and adaptively updating their posteriors with online rewards to adapt to the evolving policy distribution. Extensive experiments on instruction-following (AlpacaEval-2, Arena-Hard, MT-Bench) and reasoning (GSM8K, MMLU) benchmarks show that BayesianRouter consistently outperforms individual RMs, RM ensembling, and existing routing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AdaJudge: Adaptive Multi-Perspective Judging for Reward ModelingYongliang Miao, Yangyang Liang, Mengnan DuACL 2026 · 被引用 1 次
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM RoutingYichi Zhang, Fangzheng Xie, Shu Yang, Chong WuICLR 2026 · 被引用 1 次
- From Individual to Common: An Early Exploration of Consensus in Non-verifiable Data for Balanced Preference OptimizationShangjian Yin, Zhouxing ShiACL 2026
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He 等ICLR 2026 · 被引用 211 次
相关 Paper
- Ask a Strong LLM Judge when Your Reward Model is UncertainZhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu 等NeurIPS 2025 · 被引用 14 次
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement LearningHaozhen Zhang, Tao Feng, Jiaxuan YouNeurIPS 2025 · 被引用 81 次
- RLTHF: Targeted Human Feedback for LLM AlignmentYifei Xu, Tusher Chakraborty, Emre Kiciman, Bibek Aryal 等ICML 2025
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
- Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar 等ACL 2025 · 被引用 23 次
