Imitation Learning as State Matching via Differentiable Physics
Siwei Chen, Xiao Ma, Zhongwen Xu
Abstract
Existing imitation learning (IL) methods such as inverse reinforcement learning (IRL) usually have a double-loop training process, alternating between learning a reward function and a policy and tend to suffer long training time and high variance. In this work, we identify the benefits of differentiable physics simulators and propose a new IL method, i.e., Imitation Learning as State Matching via Differentiable Physics (ILD), which gets rid of the double-loop design and achieves significant improvements in final performance, convergence speed, and stability. The proposed ILD incorporates the differentiable physics simulator as a physics prior into its computational graph for policy learning. ILD unrolls the dynamics by sampling actions from a parameterized policy and minimizing the distance between the expert trajectory and the agent trajectory. It backpropagates the gradient into the policy via temporal physics operators, which improves the transferability to unseen environments and yields higher final performance. ILD has a single-loop structure that stabilizes and speeds up training. It dynamically selects learning objectives for each state during optimization to simplify the complex optimization landscape. Experiments show that ILD outperforms state-of-theart methods in continuous control tasks with Brax, and can be applied to deformable object manipulation tasks, generalized to unseen configurations. 1 ‡ This work is completed at the SEA AI Lab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable PhysicsZhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou et al.ICLR 2021 · 164 citations
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Scalable Differentiable Physics for Learning and ControlYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2020 · 133 citations
Related papers
- DiffMimic: Efficient Motion Mimicking with Differentiable PhysicsJiawei Ren, Cunjun Yu, Siwei Chen, Xiao Ma et al.ICLR 2023 · 3 citations
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with ToolsXingyu Lin, Zhiao Huang, Yunzhu Li, Joshua B. Tenenbaum et al.ICLR 2022 · 85 citations
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 60 citations
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 56 citations
- Reward-free World Models for Online Imitation LearningShangzhe Li, Zhiao Huang, Hao SuICML 2025
