Episodic Future Thinking Mechanism for Multi-agent Reinforcement Learning
Dongsu Lee, Minhae Kwon
Abstract
Understanding cognitive processes in multi-agent interactions is a primary goal in cognitive science. It can guide the direction of artificial intelligence (AI) research toward social decision-making in multi-agent systems, which includes uncertainty from character heterogeneity. In this paper, we introduce an episodic future thinking (EFT) mechanism for a reinforcement learning (RL) agent, inspired by cognitive processes observed in animals. To enable future thinking functionality, we first develop a multi-character policy that captures diverse characters with an ensemble of heterogeneous policies. Here, the character of an agent is defined as a different weight combination on reward components, representing distinct behavioral preferences. The future thinking agent collects observation-action trajectories of the target agents and uses the pre-trained multi-character policy to infer their characters. Once the character is inferred, the agent predicts the upcoming actions of target agents and simulates the potential future scenario. This capability allows the agent to adaptively select the optimal action, considering the predicted future scenario in multi-agent interactions. To evaluate the proposed mechanism, we consider the multi-agent autonomous driving scenario with diverse driving traits and multiple particle environments. Simulation results demonstrate that the EFT mechanism with accurate character inference leads to a higher reward than existing multi-agent solutions. We also confirm that the effect of reward improvement remains valid across societies with different levels of character diversity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5adea6b3-0af3-407d-a601-d8de734d608fCited by top-tier papers2
- Multi-agent Coordination via Flow MatchingDongsu Lee, Daehee Lee, Amy ZhangICLR 2026 · 9 citations
- : Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive EnvironmentsSangeun Park, Minhae KwonICML 2026
Builds on10
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 110 citations
- ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of MindYuanfei Wang, Fangwei Zhong, Jing Xu, Yizhou WangICLR 2022 · 103 citations
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 66 citations
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
Related papers
- Heterogeneous Skill Learning for Multi-agent TasksYuntao Liu, Yuan Li, Xinhai Xu, Yong Dou et al.NeurIPS 2022 · 33 citations
- CCP: Configurable Crowd ProfilesAndreas Panayiotou, Theodoros Kyriakou, Marilena Lemonari, Yiorgos Chrysanthou et al.SIGGRAPH 2022 · 27 citations
- Unlimited Neighborhood Interaction for Heterogeneous Trajectory PredictionFang Zheng, Le Wang, Sanping Zhou, Wei Tang et al.ICCV 2021 · 39 citations
- Trajectory Prediction in Heterogeneous Environment via Attended Ecology EmbeddingWei-Cheng Lai, Zi-Xiang Xia, Hao-Siang Lin, Lien-Feng Hsu et al.ACM MM 2020 · 27 citations
- When Is Diversity Rewarded in Cooperative Multi-Agent Learning?Michael Amir, Matteo Bettini, Amanda ProrokICLR 2026 · 4 citations
