Adversarial Dual On-Policy Distillation from Expressive Teacher
Zhenglin Wan, Jingxuan Wu, Xingrui Yu, Chubin Zhang, Mingcong Lei, Bo An, Ivor Tsang, Yang You
Abstract
Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal expert actions. Yet these methods remain offline supervised learners: the policy is trained only on expert states and receives no corrective signal on the states it actually visits. On-policy distillation (OPD) offers a natural remedy, but standard OPD assumes a strong fixed teacher, which is unavailable in demonstration-only control. We propose FA-OPD, an adversarial dual on-policy distillation method in which a Flow Matching (FM) teacher is learned from demonstrations and co-trained with a lightweight MLP student. The teacher provides two complementary signals on student rollouts. The reward channel learns an expert-likeness objective over state-action pairs and drives online exploration through long-horizon policy optimization. The action channel supplies dense local targets at student-visited states, stabilizing exploitation. FA-OPD couples them so that reward distillation enables generalization beyond point-wise demonstrations, while action distillation keeps exploration anchored near expert-like behavior. Across six robot navigation, manipulation, and locomotion benchmarks, FA-OPD beats strong baselines and shows much stronger robustness under noisy or limited demonstrations. Source code: https://github.com/vanzll/FA-OPD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86c9c082-a36f-4501-b6d9-ed3fb9844725Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Understanding Diffusion Objectives as the ELBO with Simple Data AugmentationDiederik P. Kingma, Ruiqi GaoNeurIPS 2023 · 358 citations
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk et al.ICLR 2024 · 311 citations
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim et al.ICLR 2026 · 103 citations
Related papers
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 159 citations
- Entropy-Aware On-Policy Distillation of Language ModelsWoogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei et al.ICML 2026 · 91 citations
- Variational Distillation of Diffusion Policies into Mixture of ExpertsHongyi Zhou, Denis Blessing, Ge Li, Onur Celik et al.NeurIPS 2024 · 19 citations
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 1 citation
- Diffusion Imitation from ObservationBo-Ruei Huang, Chun-Kai Yang, Chun-Mao Lai, Dai-Jie Wu et al.NeurIPS 2024 · 15 citations
