Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Shang Liu, Yu Pan, Guanting Chen, Xiaocheng Li
Abstract
Learning a reward model (RM) from human preferences has been an important component in aligning large language models (LLMs). The canonical setup of learning RMs from pairwise preference data is rooted in the classic Bradley-Terry (BT) model that accepts binary feedback, i.e., the label being either Response 1 is better than Response 2, or the opposite. Such a setup inevitably discards potentially useful samples (such as "tied" between the two responses) and loses more finegrained information (such as "slightly better"). In this paper, we propose a framework for learning RMs under ordinal feedback which generalizes the case of binary preference feedback to any arbitrary granularity. Specifically, we first identify a marginal unbiasedness condition, which generalizes the assumption of the BT model in the existing binary feedback setting. The condition validates itself via the sociological concept of the wisdom of the crowd. Under the condition, we develop a natural probability model for pairwise preference data under ordinal feedback and analyze its properties. We prove the statistical benefits of ordinal feedback in terms of reducing the Rademacher complexity compared to the case of binary feedback. The proposed learning objective and the theory also extend to hinge loss and direct policy optimization (DPO). In particular, the theoretical analysis may be of independent interest when applying to a seemingly unrelated problem of knowledge distillation to interpret the bias-variance trade-off therein. The framework also sheds light on writing guidance for human annotators. Our numerical experiments validate that fine-grained feedback leads to better reward learning for both in-distribution and out-of-distribution settings. Further experiments show that incorporating a certain proportion of samples with tied preference boosts RM learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a80fedc-cbfc-4e77-b084-47bd1299d1d4Cited by top-tier papers3
- Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal FeedbackAmirhossein Afsharrad, Ruida Zhou, Luca Viano, Sanjay Lall et al.ICLR 2026 · 4 citations
- Learning Ordinal Probabilistic Reward from PreferencesLongze Chen, Lu Wang, Renke Shan, Ze Gong et al.ICLR 2026 · 3 citations
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- RRHF: Rank Responses to Align Language Models with Human FeedbackHongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang et al.NeurIPS 2023 · 515 citations
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan et al.ICML 2024 · 447 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
Related papers
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference LearningJingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun et al.EMNLP 2025 · 2 citations
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
- Improving LLM General Preference Alignment via Optimistic Online Mirror DescentYuheng Zhang, Dian Yu, Tao Ge, Linfeng Song et al.NeurIPS 2025 · 27 citations
- Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?Paul Gölz, Nika Haghtalab, Kunhe YangNeurIPS 2025 · 29 citations
- What Does Preference Learning Recover from Pairwise Comparison Data?Rattana Pukdee, Nina Balcan, Pradeep RavikumarICML 2026 · 1 citation
