Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency
Qifan Liang, Yixiang Shan, Haipeng Liu, Zhengbang Zhu, Ting Long, Weinan Zhang, Yuan Tian
Abstract
An important challenge in multi-agent reinforcement learning is partial observability, where agents cannot access the global state of the environment during execution and can only receive observations within their field of view. To address this issue, previous works typically use the dimension-wise state. This state is obtained by applying MLP or dimension-based attention on the global state for decision-making during training and relying on a reconstructed dimensionwise state during execution. However, dimension-wise states tend to divert agent attention to specific features, neglecting potential dependencies between agents, making optimal decisions more difficult. Moreover, the inconsistency between the states used in training and execution further increases additional errors. To resolve these issues, we propose a method called Reconstruction-Guided Policy (RGP) to reconstruct the agent-wise state, which represents information of interagent relationships, as input for decision-making during both training and execution. This not only preserves the potential dependencies between agents but also ensures consistency between the states used in training and execution. We conducted extensive experiments on both discrete and continuous action environments to evaluate RGP, and the results demonstrate its superior effectiveness. Our code is public in https://github.com/Muise4/RGP4/tree/main
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99c21378-e079-4bbe-98d7-7e4d1c21f0c3Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- QPLEX: Duplex Dueling Multi-Agent Q-LearningJianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu et al.ICLR 2021 · 595 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingZekun Qi, Runpei Dong, Guofan Fan, Zheng Ge et al.ICML 2023 · 209 citations
Related papers
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 2 citations
- Consensus Learning for Cooperative Multi-Agent Reinforcement LearningZhiwei Xu, Bin Zhang, Dapeng Li, Zeren Zhang et al.AAAI 2023 · 27 citations
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte et al.ICML 2023 · 21 citations
- GRDC: A Unified Graph-Driven Framework for Role Discovery and Communication in Multi-Agent Reinforcement LearningZihong Gao, Hongjian Liang, Yuanhui Hao, Lei Hao et al.AAAI 2026
