Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning
Tianmeng Hu, Biao Luo, Ke Li
Abstract
Multi-objective reinforcement learning (MORL) seeks policies that effectively balance conflicting objectives. However, presenting many diverse policies without accounting for the decision maker’s (DM’s) preferences can overwhelm the decision-making process. On the other hand, accurately specifying preferences in advance is often unrealistic. To address these challenges, we introduce a human-in-the-loop MORL framework that interactively discovers preferred policies during optimization. Our approach proactively learns the DM’s implicit preferences in real time, requiring no a priori knowledge. Importantly, we integrate this preference learning directly into a parallel optimization framework, balancing exploration and exploitation to identify high-quality policies aligned with the DM's preferences. Evaluations on a complex quadrupedal robot simulation environment demonstrate that, with only interactions, our proposed method can identify policies aligned with human preferences, e.g., running like a dog. Further experiments on seven MuJoCo tasks and a multi-microgrid system design task against eight state-of-the-art MORL algorithms fully demonstrate the effectiveness of our proposed framework. Demonstrations and full experiments are in https://sites.google.com/view/pbmorl/home.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e8a5e3b-c064-4cba-8ad5-887037cc48b1Builds on7
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
- Understanding the automated parameter optimization on transfer learning for cross-project defect prediction: an empirical studyKe Li, Zilin Xiang, Tao Chen, Shuo Wang et al.ICSE 2020 · 54 citations
- DeepSQLi: deep semantic learning for testing SQL injectionMuyang Liu, Ke Li, Tao ChenISSTA 2020 · 47 citations
- BiLO-CPDP: Bi-Level Programming for Automated Model Discovery in Cross-Project Defect PredictionKe Li, Zilin Xiang, Tao Chen, Kay Chen TanASE 2020 · 26 citations
Related papers
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 8 citations
- Promptable Behaviors: Personalizing Multi-Objective Rewards from Human PreferencesMinyoung Hwang, Luca Weihs, Chanwoo Park, Kimin Lee et al.CVPR 2024
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 1 citation
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
