Adaptable Agent Populations via a Generative Model of Policies
Kenneth Derek, Phillip Isola
摘要
In the natural world, life has found innumerable ways to survive and often thrive. Between and even within species, each individual is in some manner unique, and this diversity lends adaptability and robustness to life. In this work, we aim to learn a space of diverse and high-reward policies on any given environment. To this end, we introduce a generative model of policies, which maps a low-dimensional latent space to an agent policy space. Our method enables learning an entire population of agent policies, without requiring the use of separate policy parameters. Just as real world populations can adapt and evolve via natural selection, our method is able to adapt to changes in our environment solely by selecting for policies in latent space. We test our generative model's capabilities in a variety of environments, including an open-ended grid-world and a two-player soccer environment. Code, visualizations, and additional experiments can be found at https://kennyderek.github.io/adap/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Diverse Conventions for Human-AI CollaborationBidipta Sarkar, Andy Shih, Dorsa SadighNeurIPS 2023 · 被引用 23 次
- Adaptive Coordination in Social Embodied RearrangementAndrew Szot, Unnat Jain, Dhruv Batra, Zsolt Kira 等ICML 2023 · 被引用 20 次
- DGPO: Discovering Multiple Strategies with Diversity-Guided Policy OptimizationWentse Chen, Shiyu Huang, Yuan Chiang, Tim Pearce 等AAAI 2024 · 被引用 9 次
- Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution TrajectoriesLi-Cheng Lan, Huan Zhang, Cho-Jui HsiehICLR 2023 · 被引用 3 次
- Improving Human-AI Coordination through Online Adversarial Training and Generative ModelsParesh R. Chaudhary, Yancheng Liang, Daphne Chen, Simon Shaolei Du 等ICLR 2026 · 被引用 2 次
相关 Paper
- Learning a subspace of policies for online adaptation in Reinforcement LearningJean-Baptiste Gaya, Laure Soulier, Ludovic DenoyerICLR 2022 · 被引用 17 次
- Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted DiffusionYongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang 等NeurIPS 2024 · 被引用 15 次
- From Parameters to Behaviors: Unsupervised Compression of the Policy SpaceDavide Tenedini, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICLR 2026 · 被引用 4 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- From Noise to Control: Parameterized Diffusion PoliciesRenhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris 等ICML 2026
