Pareto Policy Adaptation
Panagiotis Kyriakis, Jyotirmoy Deshmukh, Paul Bogdan
Abstract
We present a policy gradient method for Multi-Objective Reinforcement Learning under unknown, linear preferences. By enforcing Pareto stationarity, a first-order condition for Pareto optimality, we are able to design a simple policy gradient algorithm that approximates the Pareto front and infers the unknown preferences. Our method relies on a projected gradient descent solver that identifies common ascent directions for all objectives. Leveraging the solution of that solver, we introduce Pareto Policy Adaptation (PPA), a loss function that adapts the policy to be optimal with respect to any distribution over preferences. PPA uses implicit differentiation to back-propagate the loss gradient bypassing the operations of the projected gradient descent solver. Our approach is straightforward, easy to implement and can be used with all existing policy gradient and actor-critic methods. We evaluate our method in a series of reinforcement learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9006e3d-d248-44d1-8a40-6b5186d1937bCited by top-tier papers9
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 53 citations
- Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement LearningRuida Zhou, Tao Liu, Dileep Kalathil, P. R. Kumar et al.NeurIPS 2022 · 22 citations
- FERERO: A Flexible Framework for Preference-Guided Multi-Objective LearningLisha Chen, A F M Saif, Yanning Shen, Tianyi ChenNeurIPS 2024 · 12 citations
- COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpacePengyi Li, Hongyao Tang, Yifu Yuan, Jianye Hao et al.NeurIPS 2025 · 3 citations
- A Reward-Free Viewpoint on Multi-Objective Reinforcement LearningYing-Tu Chen, Wei Hung, Bing-Shu Wu, Zhang-Wei Hong et al.ICLR 2026 · 2 citations
Builds on2
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
Related papers
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu et al.ICML 2024 · 4 citations
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 21 citations
- Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationDebabrata Mahapatra, Vaibhav RajanICML 2020 · 182 citations
- An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement LearningQian Lin, Zongkai Liu, Danying Mo, Chao YuNeurIPS 2024 · 8 citations
- Preference Optimization on Pareto Sets: On a Theory of Multi-Objective OptimizationAbhishek Roy, Geelon So, Yian MaNeurIPS 2025 · 12 citations
