Learning Distinguishable Trajectory Representation with Contrastive Loss
Tianxu Li, Kun Zhu, Juan Li, Yang Zhang
摘要
Policy network parameter sharing is a commonly used technique in advanced deep multi-agent reinforcement learning (MARL) algorithms to improve learning effi-ciency by reducing the number of policy parameters and sharing experiences among agents. Nevertheless, agents that share the policy parameters tend to learn similar behaviors. To encourage multi-agent diversity, prior works typically maximize the mutual information between trajectories and agent identities using variational inference. However, this category of methods easily leads to inefficient exploration due to limited trajectory visitations. To resolve this limitation, inspired by the learning of pre-trained models, in this paper, we propose a novel Contrastive Trajectory Representation (CTR) method based on learning distinguishable trajectory representations to encourage multi-agent diversity. Specifically, CTR maps the trajectory of an agent into a latent trajectory representation space by an encoder and an autoregressive model. To achieve the distinguishability among trajectory representations of different agents, we introduce contrastive learning to maximize the mutual information between the trajectory representations and learnable identity representations of different agents. We implement CTR on top of QMIX and evaluate its performance in various cooperative multi-agent tasks. The empirical results demonstrate that our proposed CTR yields significant performance improvement over the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- When Is Diversity Rewarded in Cooperative Multi-Agent Learning?Michael Amir, Matteo Bettini, Amanda ProrokICLR 2026 · 被引用 4 次
- SDE-HARL: Scalable Distributed Policy Execution for Heterogeneous-Agent Reinforcement LearningToan D. Gian, Mohammad Abdi, Nathaniel D. Bastian, Francesco RestucciaAAAI 2026
- Towards Complete Multi-Agent Coordination Policy Learning via Denoising Maximum Entropy OptimizationGuanghao Li, lei yuan, Ruiqi Xue, Hengchang Zhang 等ICML 2026
它引用的顶会 Paper15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu 等ICLR 2021 · 被引用 595 次
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao 等NeurIPS 2021 · 被引用 224 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng 等NeurIPS 2021 · 被引用 133 次
相关 Paper
- Encouraging metric-aware diversity in contrastive representation spaceTianxu Li, Kun ZhuNeurIPS 2025
- Toward Efficient Multi-Agent Exploration With Trajectory Entropy MaximizationTianxu Li, Kun ZhuICLR 2025
- Contrastive Identity-Aware Learning for Multi-Agent Value DecompositionShunyu Liu, Yihe Zhou, Jie Song, Tongya Zheng 等AAAI 2023 · 被引用 43 次
- Learning Joint Behaviors with Large VariationsTianxu Li, Kun ZhuAAAI 2025 · 被引用 2 次
- Controlling Behavioral Diversity in Multi-Agent Reinforcement LearningMatteo Bettini, Ryan Kortvelesy, Amanda ProrokICML 2024 · 被引用 11 次
