On the interaction between supervision and self-play in emergent communication
Ryan Lowe, Abhinav Gupta, Jakob N. Foerster, Douwe Kiela, Joelle Pineau
Abstract
A promising approach for teaching artificial agents to use natural language involves using human-in-the-loop training. However, recent work suggests that current machine learning methods are too data inefficient to be trained in this way from scratch. In this paper, we investigate the relationship between two categories of learning signals with the ultimate goal of improving sample efficiency: imitating human language data via supervised learning, and maximizing reward in a simulated multi-agent environment via self-play (as done in emergent communication), and introduce the term supervised self-play (S2P) for algorithms using both of these signals. We find that first training agents via supervised learning on human data followed by self-play outperforms the converse, suggesting that it is not beneficial to emerge languages from scratch. We then empirically investigate various S2P schedules that begin with supervised learning in two environments: a Lewis signaling game with symbolic inputs, and an image-based referential game with natural language descriptions. Lastly, we introduce population based approaches to S2P, which further improves the performance over single-agent methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67423150-99fe-4f95-a003-88f2f8d4b49fCited by top-tier papers22
- Collaborating with Humans without Human DataDJ Strouse, Kevin R. McKee, Matt M. Botvinick, Edward Hughes et al.NeurIPS 2021 · 239 citations
- Emergent Communication of GeneralizationsJesse Mu, Noah D. GoodmanNeurIPS 2021 · 60 citations
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador et al.NeurIPS 2021 · 53 citations
- Few-shot Language Coordination by Modeling Theory of MindHao Zhu, Graham Neubig, Yonatan BiskICML 2021 · 43 citations
- Language Model Alignment with Elastic ResetMichael Noukhovitch, Samuel Lavoie, Florian Strub, Aaron C. CourvilleNeurIPS 2023 · 42 citations
Builds on1
Related papers
- Emergent Communication: Generalization and Overfitting in Lewis GamesMathieu Rita, Corentin Tallec, Paul Michel, Jean-Bastien Grill et al.NeurIPS 2022 · 41 citations
- Learning Multi-Object Positional Relationships via Emergent CommunicationYicheng Feng, Boshi An, Zongqing LuAAAI 2024 · 4 citations
- Incorporating Pragmatic Reasoning Communication into Emergent LanguageYipeng Kang, Tonghan Wang, Gerard de MeloNeurIPS 2020 · 26 citations
- Revisiting Populations in multi-agent CommunicationPaul Michel, Mathieu Rita, Kory Wallace Mathewson, Olivier Tieleman et al.ICLR 2023
- Searching for the Most Human-like Emergent LanguageBrendon Boldt, David R. MortensenEMNLP 2025
