Diffusing States and Matching Scores: A New Framework for Imitation Learning
Runzhe Wu, Yiding Chen, Gokul Swamy, Kianté Brantley, Wen Sun
摘要
Adversarial Imitation Learning is traditionally framed as a two-player zero-sum game between a learner and an adversarially chosen cost function, and can therefore be thought of as the sequential generalization of a Generative Adversarial Network (GAN). However, in recent years, diffusion models have emerged as a non-adversarial alternative to GANs that merely require training a score function via regression, yet produce generations of higher quality. In response, we investigate how to lift insights from diffusion modeling to the sequential setting. We propose diffusing states and performing score-matching along diffused states to measure the discrepancy between the expert's and learner's states. Thus, our approach only requires training score functions to predict noises via standard regression, making it significantly easier and more stable to train than adversarial methods. Theoretically, we prove first-and second-order instance-dependent bounds with linear scaling in the horizon, proving that our approach avoids the compounding errors that stymie offline approaches to imitation learning. Empirically, we show our approach outperforms both GAN-style imitation learning baselines and discriminator-free imitation learning baselines across various continuous control problems, including complex tasks like controlling humanoids to walk, sit, crawl, and navigate through obstacles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scaling Offline RL via Efficient and Expressive Shortcut ModelsNicolas A. Espinosa Dice, Yiyi Zhang, Yiding Chen, Bradley Guo 等NeurIPS 2025 · 被引用 28 次
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj 等NeurIPS 2025 · 被引用 27 次
- Fourier Features Let Agents Learn High Precision Policies with Imitation LearningBalázs Gyenes, Emiliyan Gospodinov, Jan Frieling, Enrico Krohmer 等ICML 2026 · 被引用 3 次
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 被引用 1 次
- Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation LearningAntoine Bergerault, Volkan Cevher, Negar MehrICLR 2026
它引用的顶会 Paper38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
相关 Paper
- Diffusion Imitation from ObservationBo-Ruei Huang, Chun-Kai Yang, Chun-Mao Lai, Dai-Jie Wu 等NeurIPS 2024 · 被引用 15 次
- Diffusion-Reward Adversarial Imitation LearningChun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Yu-Chiang Frank Wang 等NeurIPS 2024 · 被引用 28 次
- DiffAIL: Diffusion Adversarial Imitation LearningBingzheng Wang, Guoqiang Wu, Teng Pang, Yan Zhang 等AAAI 2024 · 被引用 24 次
- Time-series Generation by Contrastive ImitationDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2021 · 被引用 29 次
- Diffusion Based Representation LearningSarthak Mittal, Korbinian Abstreiter, Stefan Bauer, Bernhard Schölkopf 等ICML 2023 · 被引用 71 次
