Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots
Wei Hung, Bo-Kai Huang, Ping-Chun Hsieh, Xi Liu
摘要
Many real-world continuous control problems are in the dilemma of weighing the pros and cons, multi-objective reinforcement learning (MORL) serves as a generic framework of learning control policies for different preferences over objectives. However, the existing MORL methods either rely on multiple passes of explicit search for finding the Pareto front and therefore are not sample-efficient, or utilizes a shared policy network for coarse knowledge sharing among policies. To boost the sample efficiency of MORL, we propose Q-Pensieve, a policy improvement scheme that stores a collection of Q-snapshots to jointly determine the policy update direction and thereby enables data sharing at the policy level. We show that Q-Pensieve can be naturally integrated with soft policy iteration with convergence guarantee. To substantiate this concept, we propose the technique of Q replay buffer, which stores the learned Q-networks from the past iterations, and arrive at a practical actor-critic implementation. Through extensive experiments and an ablation study, we demonstrate that with much fewer samples, the proposed algorithm can outperform the benchmark MORL methods on a variety of MORL benchmark tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- The Max-Min Formulation of Multi-Objective Reinforcement Learning: From Theory to a Model-Free AlgorithmGiseung Park, Woohyeon Byeon, Seongmin Kim, Elad Havakuk 等ICML 2024 · 被引用 8 次
- COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpacePengyi Li, Hongyao Tang, Yifu Yuan, Jianye Hao 等NeurIPS 2025 · 被引用 3 次
- Constrained Multi-Objective Reinforcement Learning with Max-Min CriterionGiseung Park, Hyunyoung Nam, Woohyeon Byeon, Amir Leshem 等ICML 2026
- MODULI: Unlocking Preference Generalization via Diffusion Models for Offline Multi-Objective Reinforcement LearningYifu Yuan, Zhenrui Zheng, Zibin Dong, Jianye HaoICML 2025
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song 等ICLR 2025
它引用的顶会 Paper5
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
- Multi-objective congestion controlYiqing Ma, Han Tian, Xudong Liao, Junxue Zhang 等EuroSys 2022 · 被引用 52 次
- Pareto Policy AdaptationPanagiotis Kyriakis, Jyotirmoy Deshmukh, Paul BogdanICLR 2022 · 被引用 19 次
相关 Paper
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 被引用 8 次
- Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORLZeyu Zhao, Yueling Che, Kaichen Liu, Jian Li 等ICML 2026
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu 等ICML 2024 · 被引用 4 次
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 被引用 21 次
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 被引用 71 次
