Behaviour Distillation
Andrei Lupu, Chris Lu, Jarek Liesen, Robert Tjarko Lange, Jakob Nicolaus Foerster
Abstract
Dataset distillation aims to condense large datasets into a small number of synthetic examples that can be used as drop-in replacements when training new models. It has applications to interpretability, neural architecture search, privacy, and continual learning. Despite strong successes in supervised domains, such methods have not yet been extended to reinforcement learning, where the lack of a fixed dataset renders most distillation methods unusable. Filling the gap, we formalize behaviour distillation, a setting that aims to discover and then condense the information required for training an expert policy into a synthetic dataset of state-action pairs, without access to expert data. We then introduce Hallucinating Datasets with Evolution Strategies (HaDES), a method for behaviour distillation that can discover datasets of just four state-action pairs which, under supervised learning, train agents to competitive performance levels in continuous control tasks. We show that these datasets generalize out of distribution to training policies with a wide range of architectures and hyperparameters. We also demonstrate application to a downstream task, namely training multi-task agents in a zero-shot fashion. Beyond behaviour distillation, HaDES provides significant improvements in neuroevolution for RL over previous approaches and achieves SoTA results on one standard supervised dataset distillation task. Finally, we show that visualizing the synthetic datasets can provide human-interpretable task insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7ee3528-3e04-4387-af92-d2088e90bb31Cited by top-tier papers3
- EvIL: Evolution Strategies for Generalisable Imitation LearningSilvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh et al.ICML 2024 · 10 citations
- Offline Behavior DistillationShiye Lei, Sen Zhang, Dacheng TaoNeurIPS 2024 · 2 citations
- Algorithmic Guarantees for Distilling Supervised and Offline RL DatasetsAaryan Gupta, Rishi Saket, Aravindan RaghuveerICLR 2026
Builds on17
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 390 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
Related papers
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific ReasoningKehua Feng, Keyan Ding, Zhihui Zhu, Lei Liang et al.ICLR 2026 · 4 citations
- Powering Verifiable Learning via Automated Evolutionary Data SynthesisHe Du, Bowen Li, Aijun Yang, Siyang He et al.ACL 2026
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto et al.ICLR 2023 · 10 citations
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
- Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple BaselineHongjoon Ahn, Jinu Hyeon, Youngmin Oh, Bosun Hwang et al.ICLR 2025
