LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
Ruiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li, Zhijiang Shao
Abstract
Lexicographic multi-objective problems, which consist of multiple conflicting subtasks with explicit priorities, are common in real-world applications. Despite the advantages of Reinforcement Learning (RL) in single tasks, extending conventional RL methods to prioritized multiple objectives remains challenging. In particular, traditional Safe RL and Multi-Objective RL (MORL) methods have difficulty enforcing priority orderings efficiently. Therefore, Lexicographic Multi-Objective RL (LMORL) methods have been developed to address these challenges. However, existing LMORL methods either rely on heuristic threshold tuning with prior knowledge or are restricted to discrete domains. To overcome these limitations, we propose Lexicographically Projected Policy Gradient RL (LPPG-RL), a novel LMORL framework which leverages sequential gradient projections to identify feasible policy update directions, thereby enabling LPPG-RL broadly compatible with all policy gradient algorithms in continuous spaces. LPPG-RL reformulates the projection step as an optimization problem, and utilizes Dykstra's projection rather than generic solvers to deliver great speedups, especially for small- to medium-scale instances. In addition, LPPG-RL introduces Subproblem Exploration (SE) to prevent gradient vanishing, accelerate convergence and enhance stability. We provide theoretical guarantees for convergence and establish a lower bound on policy improvement. Finally, through extensive experiments in a 2D navigation environment, we demonstrate the effectiveness of LPPG-RL, showing that it outperforms existing state-of-the-art continuous LMORL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ceb2fe4-501b-46e3-8d16-061f685d44bcBuilds on6
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- Pareto Set Learning for Multi-Objective Reinforcement LearningErlong Liu, Yu-Chang Wu, Xiaobin Huang, Chengrui Gao et al.AAAI 2025 · 24 citations
- RLfOLD: Reinforcement Learning from Online Demonstrations in Urban Autonomous DrivingDaniel Coelho, Miguel Oliveira, Vitor SantosAAAI 2024 · 14 citations
Related papers
- Multi-objective Linear Reinforcement Learning with Lexicographic RewardsBo Xue, Dake Bu, Ji Cheng, Yuanyu Wan et al.ICML 2025
- Prioritized Soft Q-Decomposition for Lexicographic Reinforcement LearningFinn Rietz, Erik Schaffernicht, Stefan Heinrich, Johannes A. StorkICLR 2024 · 6 citations
- Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement LearningDohyeong Kim, Mineui Hong, Jeongho Park, Songhwai OhICLR 2025
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
