Countering Language Drift with Seeded Iterated Learning
Yuchen Lu, Soumye Singhal, Florian Strub, Aaron C. Courville, Olivier Pietquin
Abstract
Pretraining on human corpus and then finetuning in a simulator has become a standard pipeline for training a goal-oriented dialogue agent. Nevertheless, as soon as the agents are finetuned to maximize task completion, they suffer from the so-called language drift phenomenon: they slowly lose syntactic and semantic properties of language as they only focus on solving the task. In this paper, we propose a generic approach to counter language drift called Seeded iterated learning (SIL). We periodically refine a pretrained student agent by imitating data sampled from a newly generated teacher agent. At each time step, the teacher is created by copying the student agent, before being finetuned to maximize task completion. SIL does not require external syntactic constraint nor semantic knowledge, making it a valuable task-agnostic finetuning protocol. We evaluate SIL in a toy-setting Lewis Game, and then scale it up to the translation game with natural language. In both settings, SIL helps counter language drift as well as it improves the task completion compared to baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7107a7d8-e9d4-4aa3-a150-21def6874b3eCited by top-tier papers28
- Cones: Concept Neurons in Diffusion Models for Customized GenerationZhiheng Liu, Ruili Feng, Kai Zhu, Yifei Zhang et al.ICML 2023 · 164 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu et al.ICCV 2023 · 89 citations
- Emergent Communication at ScaleRahma Chaabouni, Florian Strub, Florent Altché, Eugene Tarassov et al.ICLR 2022 · 65 citations
- Fortuitous Forgetting in Connectionist NetworksHattie Zhou, Ankit Vani, Hugo Larochelle, Aaron C. CourvilleICLR 2022 · 50 citations
Builds on4
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- Compositional languages emerge in a neural iterated learning modelYi Ren, Shangmin Guo, Matthieu Labeau, Shay B. Cohen et al.ICLR 2020 · 111 citations
- Multi-agent Communication meets Natural Language: Synergies between Functional and Structural Language LearningAngeliki Lazaridou, Anna Potapenko, Olivier TielemanACL 2020 · 11 citations
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
Related papers
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
- GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue SystemsYoungsoo Jang, Jongmin Lee, Kee-Eung KimICLR 2022 · 45 citations
- Emergent Communication: Generalization and Overfitting in Lewis GamesMathieu Rita, Corentin Tallec, Paul Michel, Jean-Bastien Grill et al.NeurIPS 2022 · 41 citations
- Transferable Dialogue Systems and User SimulatorsBo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, Bill ByrneACL 2021
- Different Strokes for Different Folks: Investigating Appropriate Further Pre-training Approaches for Diverse Dialogue TasksYao Qiu, Jinchao Zhang, Jie ZhouEMNLP 2021 · 1 citation
