PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
Ai Jian, Jingqing Ruan, Xing Ma, Dailin Li, Weipeng Zhang, Ke Zeng, Xunliang Cai
Abstract
Reward models (RMs) are central to reinforcement learning from human feedback (RLHF), providing the critical supervision signals that align large language models (LLMs) with human preferences. Generative reward models (GRMs) provide greater interpretability than traditional scalar RMs, but they come with a critical trade-off: pairwise methods are hindered by a training-inference mismatch, while pointwise methods require expensive absolute annotations. To bridge this gap, we propose the Preference-aware Task-adaptive Reward Model (PaTaRM). Unlike prior approaches, PaTaRM enables robust pointwise training using readily available pairwise data via a novel Preference-Aware Reward (PAR) mechanism, eliminating the need for explicit rating labels. Furthermore, it incorporates a taskadaptive rubric system that dynamically generates instance-specific criteria for precise evaluation. Extensive experiments demonstrate that PaTaRM achieves an average relative improvement of 8.7% over the corresponding base models on RewardBench and RMBench across the Qwen3-8B and Qwen3-14B backbones. Crucially, when used as a reward model for downstream RLHF, it yields an average relative improvement of 13.6% over the corresponding base policies on IFEval and InfoBench, validating its effectiveness for policy alignment. Our code, data, and checkpoints are available at https://huggingface.co/AIJian/PaTaRM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Reward Modeling for Scientific Writing EvaluationFurkan Sahinuç, Subhabrata Dutta, Iryna GurevychACL 2026 · 2 citations
- IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward ModelsHaonan Song, Qingchen Xie, Huan Zhu, Feng Xiao et al.ICML 2026
Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable DomainsAnisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath et al.ICLR 2026 · 340 citations
- RM-R1: Reward Modeling as ReasoningXiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin et al.ICLR 2026 · 147 citations
- Checklists Are Better Than Reward Models For Aligning Language ModelsVijay Viswanathan, Yanchao Sun, Xiang Kong, Meng Cao et al.NeurIPS 2025 · 127 citations
Related papers
- ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment FrameworkKai Qin, Liangxin Liu, Yu Liang, Longzheng Wang et al.ACL 2026
- GRAM: A Generative Foundation Reward Model for Reward GeneralizationChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu et al.ICML 2025
- GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward ModelJingyu Zhang, Kun Yang, Ming Wen, jiawei zhao et al.ICML 2026
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun et al.NeurIPS 2025 · 17 citations
- HAF-RM: A Hybrid Alignment Framework for Reward Model TrainingShujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang et al.ACL 2025 · 4 citations
