Contextual Bilevel Reinforcement Learning for Incentive Alignment
Vinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi, Yifan Hu
Abstract
The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce Contextual Bilevel Reinforcement Learning (CB-RL), a stochastic bilevel decision-making model, where the lower level consists of solving a contextual Markov Decision Process (CMDP). CB-RL can be viewed as a Stackelberg Game where the leader and a random context beyond the leader's control together decide the setup of many MDPs that potentially multiple followers best respond to. This framework extends beyond traditional bilevel optimization and finds relevance in diverse fields such as RLHF, tax design, reward shaping, contract theory and mechanism design. We propose a stochastic Hyper Policy Gradient Descent (HPGD) algorithm to solve CB-RL, and demonstrate its convergence. Notably, HPGD uses stochastic hypergradient estimates, based on observations of the followers' trajectories. Therefore, it allows followers to use any training procedure and the leader to be agnostic of the specific algorithm, which aligns with various real-world scenarios. We further consider the setting when the leader can influence the training of followers and propose an accelerated algorithm. We empirically demonstrate the performance of our algorithm for reward shaping and tax design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23a22a8c-87b4-4548-9c95-fac37a830269Cited by top-tier papers16
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness ConditionsLiuyuan Jiang, Quan Xiao, Lisha Chen, Tianyi ChenNeurIPS 2025 · 11 citations
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningWenchang Duan, Yaoliang Yu, Jiwan He, Yi ShiNeurIPS 2025 · 11 citations
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential GameBarna Pásztor, Thomas Kleine Buening, Andreas KrauseICLR 2026 · 9 citations
- Towards Efficient Constraint Handling in Neural Solvers for Routing ProblemsJieyi Bi, Zhiguang Cao, Jianan Zhou, Wen Song et al.ICLR 2026 · 5 citations
- Conditional Gradient Methods with Standard LMO for Stochastic Simple Bilevel OptimizationKhanh-Hung Giang-Tran, Soroosh Shafiee, Nam Ho-NguyenNeurIPS 2025 · 3 citations
Builds on27
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 176 citations
- A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-MomentumPrashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai et al.NeurIPS 2021 · 175 citations
Related papers
- Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHFHan Shen, Zhuoran Yang, Tianyi ChenICML 2024 · 35 citations
- Contextual Stochastic Bilevel OptimizationYifan Hu, Jie Wang, Yao Xie, Andreas Krause et al.NeurIPS 2023 · 22 citations
- Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning AlgorithmsLiyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov et al.AAAI 2022 · 50 citations
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 137 citations
- Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg GamesKai Wang, Lily Xu, Andrew Perrault, Michael K. Reiter et al.AAAI 2022 · 29 citations
