PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, Furong Huang
Abstract
We present a novel unified bilevel optimization-based framework, PARL, formulated to address the recently highlighted critical issue of policy alignment in reinforcement learning using utility or preference-based feedback. We identify a major gap within current algorithmic designs for solving policy alignment due to a lack of precise characterization of the dependence of the alignment objective on the data generated by policy trajectories. This shortfall contributes to the sub-optimal performance observed in contemporary algorithms. Our framework addressed these concerns by explicitly parameterizing the distribution of the upper alignment objective (reward design) by the lower optimal variable (optimal policy for the designed reward). Interestingly, from an optimization perspective, our formulation leads to a new class of stochastic bilevel problems where the stochasticity at the upper objective depends upon the lower-level variable. True to our best knowledge, this work presents the first formulation of the RLHF as a bilevel optimization problem which generalizes the existing RLHF formulations and addresses the existing distribution shift issues in RLHF formulations. To demonstrate the efficacy of our formulation in resolving alignment issues in RL, we devised an algorithm named A-PARL to solve PARL problem, establishing sample complexity bounds of order . Our empirical results substantiate that the proposed PARL can address the alignment concerns in RL by showing significant improvements (up to 63% in terms of required samples) for policy alignment in large-scale environments of the Deepmind control suite and Meta world tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fce1b8e0-2c9f-4488-9444-5d1bd493db6eCited by top-tier papers12
- Contextual Bilevel Reinforcement Learning for Incentive AlignmentVinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi et al.NeurIPS 2024 · 21 citations
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li et al.ICML 2026 · 3 citations
- A Tale of Two Problems: Multi-Task Bilevel Learning Meets Equality Constrained Multi-Objective OptimizationZhiyao Zhang, Myeung Suk Oh, Zhen Qin, Jiaxiang Li et al.ICML 2026 · 1 citation
- Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward InferenceQining Zhang, Lei YingICLR 2025
- Reducing Contextual Stochastic Bilevel Optimization via Structured Function ApproximationMaxime Bouscary, Jiawei Zhang, Saurabh AminICLR 2026
Builds on24
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
Related papers
- Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningShihan Dou, Yan Liu, Enyu Zhou, Songyang Gao et al.AAAI 2025 · 2 citations
- Joint Reward and Policy Learning with Demonstrations and Human Feedback Improves AlignmentChenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li et al.ICLR 2025
- On a Connection Between Imitation Learning and RLHFTeng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen et al.ICLR 2025
- Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum ComparatorSiyuan Xu, Minghui ZhuNeurIPS 2024 · 8 citations
- Unifying Stable Optimization and Reference Regularization in RLHFLi He, Qiang Qu, He Zhao, Stephen Wan et al.ICLR 2026 · 6 citations
