Group Robust Preference Optimization in Reward-free RLHF
Shyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou-Ammar, Ilija Bogunovic
Abstract
Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional RLHF approaches adopt a"one-size-fits-all"approach, i.e., they indiscriminately assume and optimize a single preference model, thus not being robust to unique characteristics and needs of the various groups. To address this limitation, we propose a novel Group Robust Preference Optimization (GRPO) method to align LLMs to individual groups' preferences robustly. Our approach builds upon reward-free direct preference optimization methods, but unlike previous approaches, it seeks a robust policy which maximizes the worst-case group performance. To achieve this, GRPO adaptively and sequentially weights the importance of different groups, prioritizing groups with worse cumulative loss. We theoretically study the feasibility of GRPO and analyze its convergence for the log-linear policy class. By fine-tuning LLMs with GRPO using diverse group-based global opinion data, we significantly improved performance for the worst-performing groups, reduced loss imbalances across groups, and improved probability accuracies compared to non-robust baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dede7be0-84fd-41b5-a54b-cac9e002b243Cited by top-tier papers18
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making AbilitiesThomas Schmied, Jörg Bornschein, Jordi Grau-Moya, Markus Wulfmeier et al.ICLR 2026 · 41 citations
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and DiagnosisYuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue et al.CVPR 2026 · 11 citations
- UniGame: Turning a Unified Multimodal Model Into Its Own AdversaryZhaolong Su, Wang Lu, Hao Chen, Sharon Li et al.CVPR 2026 · 11 citations
- Strategyproof Reinforcement Learning from Human FeedbackThomas Kleine Buening, Jiarui Gan, Debmalya Mandal, Marta KwiatkowskaNeurIPS 2025 · 10 citations
- Multi-Task GRPO: Reliable LLM Reasoning Across TasksShyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon et al.ICML 2026 · 8 citations
Builds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- Group Preference Optimization: Few-Shot Alignment of Large Language ModelsSiyan Zhao, John Dang, Aditya GroverICLR 2024 · 54 citations
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue et al.ACL 2025 · 13 citations
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu et al.NeurIPS 2024 · 119 citations
- MaxMin-RLHF: Alignment with Diverse Human PreferencesSouradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel et al.ICML 2024 · 104 citations
- Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM AlignmentRuoxi Cheng, Haoxuan Ma, Weixin Wang, Ranjie Duan et al.ICLR 2026 · 23 citations
