Learning to Guide and to be Guided in the Architect-Builder Problem
Paul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier, Christopher Pal, Pierre-Yves Oudeyer
Abstract
We are interested in interactive agents that learn to coordinate, namely, a builder -which performs actions but ignores the goal of the task, i.e. has no access to rewards -and an architect which guides the builder towards the goal of the task. We define and explore a formal setting where artificial agents are equipped with mechanisms that allow them to simultaneously learn a task while at the same time evolving a shared communication protocol. Ideally, such learning should only rely on high-level communication priors and be able to handle a large variety of tasks and meanings while deriving communication protocols that can be reused across tasks. The field of Experimental Semiotics has shown the extent of human proficiency at learning from a priori unknown instructions meanings. Therefore, we take inspiration from it and present the Architect-Builder Problem (ABP): an asymmetrical setting in which an architect must learn to guide a builder towards constructing a specific structure. The architect knows the target structure but cannot act in the environment and can only send arbitrary messages to the builder. The builder on the other hand can act in the environment, but receives no rewards nor has any knowledge about the task, and must learn to solve it relying only on the messages sent by the architect. Crucially, the meaning of messages is initially not defined nor shared between the agents but must be negotiated throughout learning. Under these constraints, we propose Architect-Builder Iterated Guiding (ABIG), a solution to the Architect-Builder Problem where the architect leverages a learned model of the builder to guide it while the builder uses self-imitation learning to reinforce its guided behavior. To palliate to the non-stationarity induced by the two agents concurrently learning, ABIG structures the sequence of interactions between the agents into interaction frames. We analyze the key learning mechanisms of ABIG and test it in a 2-dimensional instantiation of the ABP where tasks involve grasping cubes, placing them at a given location, or building various shapes. In this environment, ABIG results in a low-level, high-frequency, guiding communication protocol that not only enables an architect-builder pair to solve the task at hand, but that can also generalize to unseen tasks. * Equal contribution. † Work conducted while at Inria.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d7f10bb-4f9f-4003-b61e-ea3d504cf676Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux et al.NeurIPS 2020 · 139 citations
- Emergent Social Learning via Multi-agent Reinforcement LearningKamal Ndousse, Douglas Eck, Sergey Levine, Natasha JaquesICML 2021 · 61 citations
- Learning to Interactively Learn and AssistMark Woodward, Chelsea Finn, Karol HausmanAAAI 2020 · 37 citations
- Interactive Learning from Activity DescriptionKhanh Nguyen, Dipendra Misra, Robert E. Schapire, Miroslav Dudík et al.ICML 2021 · 36 citations
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement LearningJulien Roy, Paul Barde, Félix G. Harvey, Derek Nowrouzezahrai et al.NeurIPS 2020 · 25 citations
Related papers
- Learning Intuitive Policies Using Action FeaturesMingwei Ma, Jizhou Liu, Samuel Sokota, Max Kleiman-Weiner et al.ICML 2023 · 4 citations
- Compositional languages emerge in a neural iterated learning modelYi Ren, Shangmin Guo, Matthieu Labeau, Shay B. Cohen et al.ICLR 2020 · 111 citations
- Evolving AgentsLeonardo RanaldiACL 2026 · 227 citations
- Bridging semantics and pragmatics in information-theoretic emergent communicationEleonora Gualdoni, Mycal Tucker, Roger Levy, Noga ZaslavskyNeurIPS 2024 · 7 citations
- Learning to execute instructions in a Minecraft dialoguePrashant Jayannavar, Anjali Narayan-Chen, Julia HockenmaierACL 2020 · 22 citations
