Prioritized Soft Q-Decomposition for Lexicographic Reinforcement Learning
Finn Rietz, Erik Schaffernicht, Stefan Heinrich, Johannes A. Stork
摘要
Reinforcement learning (RL) for complex tasks remains a challenge, primarily due to the difficulties of engineering scalar reward functions and the inherent inefficiency of training models from scratch. Instead, it would be better to specify complex tasks in terms of elementary subtasks and to reuse subtask solutions whenever possible. In this work, we address continuous space lexicographic multi-objective RL problems, consisting of prioritized subtasks, which are notoriously difficult to solve. We show that these can be scalarized with a subtask transformation and then solved incrementally using value decomposition. Exploiting this insight, we propose prioritized soft Q-decomposition (PSQD), a novel algorithm for learning and adapting subtask solutions under lexicographic priorities in continuous state-action spaces. PSQD offers the ability to reuse previously learned subtask solutions in a zero-shot composition, followed by an adaptation step. Its ability to use retained subtask training data for offline learning eliminates the need for new environment interaction during adaptation. We demonstrate the efficacy of our approach by presenting successful learning, reuse, and adaptation results for both low- and high-dimensional simulated robot control tasks, as well as offline learning results. In contrast to baseline approaches, PSQD does not trade off between conflicting subtasks or priority constraints and satisfies subtask priorities during learning. PSQD provides an intuitive framework for tackling complex RL problems, offering insights into the inner workings of the subtask composition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- APC-RL: Exceeding data-driven behavior priors with adaptive policy compositionFinn Rietz, Pedro Zuidberg Dos Martires, Johannes A. StorkICLR 2026
- LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem ExplorationRuiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li 等AAAI 2026
相关 Paper
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
- In-Context Compositional Q-Learning for Offline Reinforcement LearningQiushui Xu, Yuhao Huang, Yushu Jiang, Wenliang Zheng 等ICLR 2026
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement LearningJinmin He, Kai Li, Yifan Zang, Haobo Fu 等ICML 2025
- Value Function Decomposition for Iterative Design of Reinforcement Learning AgentsJames MacGlashan, Evan Archer, Alisa Devlic, Takuma Seno 等NeurIPS 2022 · 被引用 12 次
- Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement LearningYonghyeon Jo, Sunwoo Lee, Seungyul HanICLR 2026 · 被引用 5 次
