DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-Learning
Wei-Fang Sun, Cheng-Kuang Lee, Chun-Yi Lee
摘要
In fully cooperative multi-agent reinforcement learning (MARL) settings, the environments are highly stochastic due to the partial observability of each agent and the continuously changing policies of the other agents. To address the above issues, we integrate distributional RL and value function factorization methods by proposing a Distributional Value Function Factorization (DFAC) framework to generalize expected value function factorization methods to their DFAC variants. DFAC extends the individual utility functions from deterministic variables to random variables, and models the quantile function of the total return as a quantile mixture. To validate DFAC, we demonstrate DFAC's ability to factorize a simple two-step matrix game with stochastic rewards and perform experiments on all Super Hard tasks of StarCraft Multi-Agent Challenge, showing that DFAC is able to outperform expected value function factorization baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Heterogeneous Agent Q-weighted Policy OptimizationBor-Jiun Lin, Chun-Yi LeeICLR 2026 · 被引用 102 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value FactorizationJianhao Wang, Zhizhou Ren, Beining Han, Jianing Ye 等NeurIPS 2021 · 被引用 50 次
- Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement LearningYitian Hong, Yaochu Jin, Yang TangNeurIPS 2022 · 被引用 40 次
- ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Mengwei Qiu, Jun Liu, Weiquan Liu 等NeurIPS 2022 · 被引用 35 次
它引用的顶会 Paper1
相关 Paper
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningJianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun KimNeurIPS 2022 · 被引用 50 次
- Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value FactorizationChengrui Qu, Christopher Yeh, Kishan Panaganti, Eric Mazumdar 等ICLR 2026
- Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement LearningChang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou 等ICLR 2026
- PAC: Assisted Value Factorization with Counterfactual Predictions in Multi-Agent Reinforcement LearningHanhan Zhou, Tian Lan, Vaneet AggarwalNeurIPS 2022 · 被引用 47 次
- More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy FactorizationJiangxing Wang, Deheng Ye, Zongqing LuICLR 2023 · 被引用 5 次
