Contextual Bilevel Reinforcement Learning for Incentive Alignment
Vinzenz Thoma, Barna Pásztor, Andreas Krause, Giorgia Ramponi, Yifan Hu
摘要
The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce Contextual Bilevel Reinforcement Learning (CB-RL), a stochastic bilevel decision-making model, where the lower level consists of solving a contextual Markov Decision Process (CMDP). CB-RL can be viewed as a Stackelberg Game where the leader and a random context beyond the leader's control together decide the setup of many MDPs that potentially multiple followers best respond to. This framework extends beyond traditional bilevel optimization and finds relevance in diverse fields such as RLHF, tax design, reward shaping, contract theory and mechanism design. We propose a stochastic Hyper Policy Gradient Descent (HPGD) algorithm to solve CB-RL, and demonstrate its convergence. Notably, HPGD uses stochastic hypergradient estimates, based on observations of the followers' trajectories. Therefore, it allows followers to use any training procedure and the leader to be agnostic of the specific algorithm, which aligns with various real-world scenarios. We further consider the setting when the leader can influence the training of followers and propose an accelerated algorithm. We empirically demonstrate the performance of our algorithm for reward shaping and tax design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness ConditionsLiuyuan Jiang, Quan Xiao, Lisha Chen, Tianyi ChenNeurIPS 2025 · 被引用 11 次
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningWenchang Duan, Yaoliang Yu, Jiwan He, Yi ShiNeurIPS 2025 · 被引用 11 次
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential GameBarna Pásztor, Thomas Kleine Buening, Andreas KrauseICLR 2026 · 被引用 9 次
- Towards Efficient Constraint Handling in Neural Solvers for Routing ProblemsJieyi Bi, Zhiguang Cao, Jianan Zhou, Wen Song 等ICLR 2026 · 被引用 5 次
- Conditional Gradient Methods with Standard LMO for Stochastic Simple Bilevel OptimizationKhanh-Hung Giang-Tran, Soroosh Shafiee, Nam Ho-NguyenNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper27
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel ProblemsTianyi Chen, Yuejiao Sun, Wotao YinNeurIPS 2021 · 被引用 176 次
- A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-MomentumPrashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai 等NeurIPS 2021 · 被引用 175 次
相关 Paper
- Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHFHan Shen, Zhuoran Yang, Tianyi ChenICML 2024 · 被引用 35 次
- Contextual Stochastic Bilevel OptimizationYifan Hu, Jie Wang, Yao Xie, Andreas Krause 等NeurIPS 2023 · 被引用 22 次
- Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning AlgorithmsLiyuan Zheng, Tanner Fiez, Zane Alumbaugh, Benjamin Chasnov 等AAAI 2022 · 被引用 50 次
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 被引用 137 次
- Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg GamesKai Wang, Lily Xu, Andrew Perrault, Michael K. Reiter 等AAAI 2022 · 被引用 29 次
