Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement Learning
Tianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin Li
Abstract
Designing accurate reward functions for reinforcement learning (RL) has long been challenging. Preference-based RL (PbRL) offers a promising approach by using human preferences to train agents, eliminating the need for manual reward design. While successful in single-agent tasks, extending PbRL to complex multi-agent scenarios is nontrivial. Existing PbRL methods lack the capacity to comprehensively capture both temporal and cooperative aspects, leading to inadequate reward functions. This work introduces an advanced multi-agent preference learning framework that effectively addresses these limitations. Based on a cascaded Transformer architecture, our approach captures both temporal and cooperative dependencies, alleviating issues related to reward uniformity and intricate interactions among agents. Experimental results demonstrate substantial performance improvements in multi-agent cooperative tasks, and the reconstructed reward function closely resembles expertdefined reward functions. The source code is available at https://github.com/catezi/MAPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7980a74-37d8-4f34-81fb-17ac09d8f25aCited by top-tier papers3
- Preference Alignment with Flow MatchingMinu Kim, Yongsik Lee, Sehyeok Kang, Jihwan Oh et al.NeurIPS 2024 · 15 citations
- Grounded Answers for Multi-agent Decision-making Problem through Generative World ModelZeyang Liu, Xinrui Yang, Shiguang Sun, Long Qian et al.NeurIPS 2024 · 10 citations
- Leveraging Sub-Optimal Data for Human-in-the-Loop Reinforcement LearningCalarina Muslimani, Matthew E. TaylorICLR 2025
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
Related papers
- Preference Transformer: Modeling Human Preferences using Transformers for RLChangyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee et al.ICLR 2023 · 4 citations
- O-MAPL: Offline Multi-agent Preference LearningThe Viet Bui, Tien Mai, Thanh Hong NguyenICML 2025
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningRunze Liu, Fengshuo Bai, Yali Du, Yaodong YangNeurIPS 2022 · 72 citations
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
