Can You Rely on Synthetic Labellers in Preference-Based Reinforcement Learning? It's Complicated
Katherine Metcalf, Miguel Sarabia, Masha Fedzechkina, Barry-John Theobald
Abstract
Preference-based Reinforcement Learning (PbRL) enables non-experts to train Reinforcement Learning models using preference feedback. However, the effort required to collect preference labels from real humans means that PbRL research primarily relies on synthetic labellers. We validate the most common synthetic labelling strategy by comparing against labels collected from a crowd of humans on three Deep Mind Control (DMC) suite tasks: stand, walk, and run. We find that:
(1) the synthetic labels are a good proxy for real humans under some circumstances, (2) strong preference label agreement between human and synthetic labels is not necessary for similar policy performance, (3) policy performance is higher at the start of training from human feedback and is higher at the end of training from synthetic feedback, and (4) training on only examples with high levels of inter-annotator agreement does not meaningfully improve policy performance. Our results justify the use of synthetic labellers to develop and ablate PbRL methods, and provide insight into how human labelling changes over the course of policy training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21aa504f-9b6e-40b8-b545-01a258084515Cited by top-tier papers1
Ask how each one uses itBuilds on5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee et al.ICLR 2022 · 116 citations
- Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningRunze Liu, Fengshuo Bai, Yali Du, Yaodong YangNeurIPS 2022 · 72 citations
- Reward Uncertainty for Exploration in Preference-based Reinforcement LearningXinran Liang, Katherine Shu, Kimin Lee, Pieter AbbeelICLR 2022
Related papers
- Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement LearningBrett Barkley, David Fridovich-KeilICML 2025
- Preference-based Reinforcement Learning with Finite-Time GuaranteesYichong Xu, Ruosong Wang, Lin F. Yang, Aarti Singh et al.NeurIPS 2020 · 82 citations
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu et al.ICLR 2025
- Preference Elicitation for Offline Reinforcement LearningAlizée Pace, Bernhard Schölkopf, Gunnar Rätsch, Giorgia RamponiICLR 2025
