Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots
Wei Hung, Bo-Kai Huang, Ping-Chun Hsieh, Xi Liu
Abstract
Many real-world continuous control problems are in the dilemma of weighing the pros and cons, multi-objective reinforcement learning (MORL) serves as a generic framework of learning control policies for different preferences over objectives. However, the existing MORL methods either rely on multiple passes of explicit search for finding the Pareto front and therefore are not sample-efficient, or utilizes a shared policy network for coarse knowledge sharing among policies. To boost the sample efficiency of MORL, we propose Q-Pensieve, a policy improvement scheme that stores a collection of Q-snapshots to jointly determine the policy update direction and thereby enables data sharing at the policy level. We show that Q-Pensieve can be naturally integrated with soft policy iteration with convergence guarantee. To substantiate this concept, we propose the technique of Q replay buffer, which stores the learned Q-networks from the past iterations, and arrive at a practical actor-critic implementation. Through extensive experiments and an ablation study, we demonstrate that with much fewer samples, the proposed algorithm can outperform the benchmark MORL methods on a variety of MORL benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2caff2d-1fac-4edc-9d9a-bdd9591bfeefCited by top-tier papers8
- The Max-Min Formulation of Multi-Objective Reinforcement Learning: From Theory to a Model-Free AlgorithmGiseung Park, Woohyeon Byeon, Seongmin Kim, Elad Havakuk et al.ICML 2024 · 8 citations
- COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpacePengyi Li, Hongyao Tang, Yifu Yuan, Jianye Hao et al.NeurIPS 2025 · 3 citations
- Constrained Multi-Objective Reinforcement Learning with Max-Min CriterionGiseung Park, Hyunyoung Nam, Woohyeon Byeon, Amir Leshem et al.ICML 2026
- MODULI: Unlocking Preference Generalization via Diffusion Models for Offline Multi-Objective Reinforcement LearningYifu Yuan, Zhenrui Zheng, Zibin Dong, Jianye HaoICML 2025
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song et al.ICLR 2025
Builds on5
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
- Multi-objective congestion controlYiqing Ma, Han Tian, Xudong Liao, Junxue Zhang et al.EuroSys 2022 · 52 citations
- Pareto Policy AdaptationPanagiotis Kyriakis, Jyotirmoy Deshmukh, Paul BogdanICLR 2022 · 19 citations
Related papers
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 8 citations
- Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORLZeyu Zhao, Yueling Che, Kaichen Liu, Jian Li et al.ICML 2026
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu et al.ICML 2024 · 4 citations
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 21 citations
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 71 citations
