Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
Jianxiang Zang, Yongda Wei, Ruxue Bai, Shiyu Jiang, Nijia Mo, Binhong Li, Qiang Sun, Hui Liu
Abstract
Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods focus solely on preference perception accuracies in specific scenarios, obscuring the critical vulnerabilities of RMs in real-world scenarios. We identify that the true challenge lies in assessing a novel dimension: Suitability, defined as conditional reliability under specific realworld perturbations. To this end, we introduce Reward Auditor, a hypothesis-testing framework specifically designed for RM suitability inference. Rather than answering "How accurate is the RM's preference perception for given samples?", it employs scientific auditing to answer: "Can we infer that RMs exhibit systematic vulnerabilities in specific real-world scenarios?". Under real-world perturbed scenarios, Reward Auditor quantifies statistical significance and effect size by auditing distribution degradation of RM preference perception confidence. This enables inference of both the certainty and severity of RM vulnerabilities across real-world scenarios, thereby laying a solid foundation for building next-generation LLM alignment systems that are verifiably safe, more robust, and trustworthy. Codes available: https:// github.com/hggzjx/RewardAuditor .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bdf56a6-972d-4177-b91e-452cb963df16Cited by top-tier papers4
- BOOSTAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward ModelsYuanhao Li, Hongbo Wang, Xiaotang Shang, Xunzhu Tang et al.ICML 2026 · 2 citations
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen et al.ICML 2026
- In-Context Learning as Rate–Distortion OptimizationJiayu Zhang, Changbang Li, Canran XiaoICML 2026
- From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor DefenseBinyan Xu, Fan YANG, Xilin Dai, Di Tang et al.ICML 2026
Builds on5
- Provably Robust DPO: Aligning Language Models with Noisy FeedbackSayak Ray Chowdhury, Anush Kini, Nagarajan NatarajanICML 2024 · 118 citations
- ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry AreaJunxian Li, Di Zhang, Xunzhi Wang, Zeying Hao et al.AAAI 2025 · 71 citations
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo et al.CVPR 2026 · 61 citations
- IAG: Input-aware Backdoor Attack on VLM-based Visual GroundingJunxian Li, Beining Xu, Simin Chen, Jiatong Li et al.CVPR 2026 · 13 citations
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 2 citations
Related papers
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM ObjectivesMatthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna et al.ICLR 2026 · 1 citation
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsChenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu et al.AAAI 2026 · 1 citation
- ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment FrameworkKai Qin, Liangxin Liu, Yu Liang, Longzheng Wang et al.ACL 2026
- RMB: Comprehensively benchmarking reward models in LLM alignmentEnyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi et al.ICLR 2025
- Interpreting Language Reward Models via Contrastive ExplanationsJunqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.ICLR 2025
