Learning to Guide and to be Guided in the Architect-Builder Problem
Paul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier, Christopher Pal, Pierre-Yves Oudeyer
摘要
We are interested in interactive agents that learn to coordinate, namely, a builder -which performs actions but ignores the goal of the task, i.e. has no access to rewards -and an architect which guides the builder towards the goal of the task. We define and explore a formal setting where artificial agents are equipped with mechanisms that allow them to simultaneously learn a task while at the same time evolving a shared communication protocol. Ideally, such learning should only rely on high-level communication priors and be able to handle a large variety of tasks and meanings while deriving communication protocols that can be reused across tasks. The field of Experimental Semiotics has shown the extent of human proficiency at learning from a priori unknown instructions meanings. Therefore, we take inspiration from it and present the Architect-Builder Problem (ABP): an asymmetrical setting in which an architect must learn to guide a builder towards constructing a specific structure. The architect knows the target structure but cannot act in the environment and can only send arbitrary messages to the builder. The builder on the other hand can act in the environment, but receives no rewards nor has any knowledge about the task, and must learn to solve it relying only on the messages sent by the architect. Crucially, the meaning of messages is initially not defined nor shared between the agents but must be negotiated throughout learning. Under these constraints, we propose Architect-Builder Iterated Guiding (ABIG), a solution to the Architect-Builder Problem where the architect leverages a learned model of the builder to guide it while the builder uses self-imitation learning to reinforce its guided behavior. To palliate to the non-stationarity induced by the two agents concurrently learning, ABIG structures the sequence of interactions between the agents into interaction frames. We analyze the key learning mechanisms of ABIG and test it in a 2-dimensional instantiation of the ABP where tasks involve grasping cubes, placing them at a given location, or building various shapes. In this environment, ABIG results in a low-level, high-frequency, guiding communication protocol that not only enables an architect-builder pair to solve the task at hand, but that can also generalize to unseen tasks. * Equal contribution. † Work conducted while at Inria.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 等NeurIPS 2020 · 被引用 139 次
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 被引用 61 次
- Learning to Interactively Learn and AssistMark Woodward, Chelsea Finn, Karol HausmanAAAI 2020 · 被引用 37 次
- Interactive Learning from Activity DescriptionKhanh Nguyen, Dipendra Misra, Robert E. Schapire, Miroslav Dudík 等ICML 2021 · 被引用 36 次
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai 等NeurIPS 2020 · 被引用 25 次
相关 Paper
- Learning Intuitive Policies Using Action FeaturesMingwei Ma, Jizhou Liu, Samuel Sokota, Max Kleiman-Weiner 等ICML 2023 · 被引用 4 次
- Compositional languages emerge in a neural iterated learning modelYi Ren, Shangmin Guo, Matthieu Labeau, Shay B. Cohen 等ICLR 2020 · 被引用 111 次
- Evolving AgentsLeonardo RanaldiACL 2026 · 被引用 227 次
- Bridging semantics and pragmatics in information-theoretic emergent communicationEleonora Gualdoni, Mycal Tucker, Roger Levy, Noga ZaslavskyNeurIPS 2024 · 被引用 7 次
- Learning to execute instructions in a Minecraft dialoguePrashant Jayannavar, Anjali Narayan-Chen, Julia HockenmaierACL 2020 · 被引用 22 次
