PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning Algorithm
Toygun Basaklar, Suat Gumussoy, Ümit Y. Ogras
摘要
Multi-objective reinforcement learning (MORL) approaches have emerged to tackle many real-world problems with multiple conflicting objectives by maximizing a joint objective function weighted by a preference vector. These approaches find fixed customized policies corresponding to preference vectors specified during training. However, the design constraints and objectives typically change dynamically in real-life scenarios. Furthermore, storing a policy for each potential preference is not scalable. Hence, obtaining a set of Pareto front solutions for the entire preference space in a given domain with a single training is critical. To this end, we propose a novel MORL algorithm that trains a single universal network to cover the entire preference space scalable to continuous robotic tasks. The proposed approach, Preference-Driven MORL (PD-MORL), utilizes the preferences as guidance to update the network parameters. It also employs a novel parallelization approach to increase sample efficiency. We show that PD-MORL achieves up to 25% larger hypervolume for challenging continuous control tasks and uses an order of magnitude fewer trainable parameters compared to prior approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Decoding-Time Language Model Alignment with Multiple ObjectivesRuizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu 等NeurIPS 2024 · 被引用 111 次
- Panacea: Pareto Alignment via Preference Adaptation for LLMsYifan Zhong, Chengdong Ma, Xiaoyuan Zhang, Ziran Yang 等NeurIPS 2024 · 被引用 89 次
- Hypervolume Maximization: A Geometric View of Pareto Set LearningXiaoyuan Zhang, Xi Lin, Bo Xue, Yifan Chen 等NeurIPS 2023 · 被引用 40 次
- Pareto Set Learning for Multi-Objective Reinforcement LearningErlong Liu, Yu-Chang Wu, Xiaobin Huang, Chengrui Gao 等AAAI 2025 · 被引用 24 次
- A Hierarchical Adaptive Multi-Task Reinforcement Learning Framework for Multiplier Circuit DesignZhihai Wang, Jie Wang, Dongsheng Zuo, Yunjie Ji 等ICML 2024 · 被引用 16 次
它引用的顶会 Paper3
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 被引用 189 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
相关 Paper
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song 等ICLR 2025
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- Distributional Pareto-Optimal Multi-Objective Reinforcement LearningXin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian 等NeurIPS 2023 · 被引用 46 次
- Population-Free Pareto Tracking for Sample-Efficient Multi-Policy MORLZeyu Zhao, Yueling Che, Kaichen Liu, Jian Li 等ICML 2026
- PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao LuoAAAI 2024 · 被引用 7 次
