Behaviour Distillation
Andrei Lupu, Chris Lu, Jarek Liesen, Robert Tjarko Lange, Jakob Nicolaus Foerster
摘要
Dataset distillation aims to condense large datasets into a small number of synthetic examples that can be used as drop-in replacements when training new models. It has applications to interpretability, neural architecture search, privacy, and continual learning. Despite strong successes in supervised domains, such methods have not yet been extended to reinforcement learning, where the lack of a fixed dataset renders most distillation methods unusable. Filling the gap, we formalize behaviour distillation, a setting that aims to discover and then condense the information required for training an expert policy into a synthetic dataset of state-action pairs, without access to expert data. We then introduce Hallucinating Datasets with Evolution Strategies (HaDES), a method for behaviour distillation that can discover datasets of just four state-action pairs which, under supervised learning, train agents to competitive performance levels in continuous control tasks. We show that these datasets generalize out of distribution to training policies with a wide range of architectures and hyperparameters. We also demonstrate application to a downstream task, namely training multi-task agents in a zero-shot fashion. Beyond behaviour distillation, HaDES provides significant improvements in neuroevolution for RL over previous approaches and achieves SoTA results on one standard supervised dataset distillation task. Finally, we show that visualizing the synthetic datasets can provide human-interpretable task insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EvIL: Evolution Strategies for Generalisable Imitation LearningSilvia Sapora, Gokul Swamy, Chris Lu, Yee Whye Teh 等ICML 2024 · 被引用 10 次
- Offline Behavior DistillationShiye Lei, Sen Zhang, Dacheng TaoNeurIPS 2024 · 被引用 2 次
- Algorithmic Guarantees for Distilling Supervised and Offline RL DatasetsAaryan Gupta, Rishi Saket, Aravindan RaghuveerICLR 2026
它引用的顶会 Paper17
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Dataset Condensation with Differentiable Siamese AugmentationBo Zhao, Hakan BilenICML 2021 · 被引用 390 次
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 被引用 307 次
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu 等ICLR 2022 · 被引用 203 次
相关 Paper
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific ReasoningKehua Feng, Keyan Ding, Zhihui Zhu, Lei Liang 等ICLR 2026 · 被引用 4 次
- Powering Verifiable Learning via Automated Evolutionary Data SynthesisHe Du, Bowen Li, Aijun Yang, Siyang He 等ACL 2026
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto 等ICLR 2023 · 被引用 10 次
- Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksSiddharth Joshi, Jiayi Ni, Baharan MirzasoleimanICLR 2025
- Prevalence of Negative Transfer in Continual Reinforcement Learning: Analyses and a Simple BaselineHongjoon Ahn, Jinu Hyeon, Youngmin Oh, Bosun Hwang 等ICLR 2025
