Ad Hoc Teamwork via Offline Goal-Based Decision Transformers
Xinzhi Zhang, Hohei Chan, Deheng Ye, Yi Cai, Mengchen Zhao
摘要
The ability of agents to collaborate with previously unknown teammates on the fly, known as ad hoc teamwork (AHT), is crucial in many realworld applications. Existing approaches to AHT require online interactions with the environment and some carefully designed teammates. However, these prerequisites can be infeasible in practice. In this work, we extend the AHT problem to the offline setting, where the policy of the ego agent is directly learned from a multi-agent interaction dataset. We propose a hierarchical sequence modeling framework called TAGET that addresses critical challenges in the offline setting, including limited data, partial observability and online adaptation. The core idea of TAGET is to dynamically predict teammate-aware rewards-togo and sub-goals, so that the ego agent can adapt to the changes of teammates' behaviors in real time. Extensive experimental results show that TAGET significantly outperforms existing solutions to AHT in the offline setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Online Decision TransformerQinqing Zheng, Amy Zhang, Aditya GroverICML 2022 · 被引用 256 次
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen 等NeurIPS 2020 · 被引用 225 次
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu 等ICML 2022 · 被引用 194 次
相关 Paper
- Online Ad Hoc Teamwork under Partial ObservabilityPengjie Gu, Mengchen Zhao, Jianye Hao, Bo AnICLR 2022 · 被引用 35 次
- Opponent Modeling based on Subgoal InferenceXiaopeng Yu, Jiechuan Jiang, Zongqing LuNeurIPS 2024 · 被引用 7 次
- AATEAM: Achieving the Ad Hoc Teamwork by Employing the Attention MechanismShuo Chen, Ewa Andrejczuk, Zhiguang Cao, Jie ZhangAAAI 2020 · 被引用 54 次
- HTAC: Hierarchical Task-Aware Composition for Continual Offline Reinforcement LearningQiyang Zhou, Xu Ruihang, Peng Wang, Wenjie Lu 等ICML 2026
- Learning Generalizable Skills from Offline Multi-Task Data for Multi-Agent CooperationSicong Liu, Yang Shu, Chenjuan Guo, Bin YangICLR 2025
