A Coupled Flow Approach to Imitation Learning
Gideon Joseph Freund, Elad Sarafian, Sarit Kraus
Abstract
In reinforcement learning and imitation learning, an object of central importance is the state distribution induced by the policy. It plays a crucial role in the policy gradient theorem, and references to it--along with the related state-action distribution--can be found all across the literature. Despite its importance, the state distribution is mostly discussed indirectly and theoretically, rather than being modeled explicitly. The reason being an absence of appropriate density estimation tools. In this work, we investigate applications of a normalizing flow-based model for the aforementioned distributions. In particular, we use a pair of flows coupled through the optimality point of the Donsker-Varadhan representation of the Kullback-Leibler (KL) divergence, for distribution matching based imitation learning. Our algorithm, Coupled Flow Imitation Learning (CFIL), achieves state-of-the-art performance on benchmark tasks with a single expert trajectory and extends naturally to a variety of other settings, including the subsampled and state-only regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85422a54-9137-4d87-be64-b3a97be061edCited by top-tier papers8
- AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based PoliciesXixi Hu, Qiang Liu, Xingchao Liu, Bo LiuNeurIPS 2024 · 73 citations
- Diffusion-Reward Adversarial Imitation LearningChun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Yu-Chiang Frank Wang et al.NeurIPS 2024 · 28 citations
- DiffAIL: Diffusion Adversarial Imitation LearningBingzheng Wang, Guoqiang Wu, Teng Pang, Yan Zhang et al.AAAI 2024 · 24 citations
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu et al.ICLR 2024 · 17 citations
- A Simple Solution for Offline Imitation from Observations and Examples with Possibly Incomplete TrajectoriesKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2023 · 7 citations
Builds on9
- Why Normalizing Flows Fail to Detect Out-of-Distribution DataPolina Kirichenko, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 370 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- Variational Policy Gradient Method for Reinforcement Learning with General UtilitiesJunyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvári et al.NeurIPS 2020 · 170 citations
- Off-Policy Imitation Learning from ObservationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouNeurIPS 2020 · 102 citations
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 96 citations
Related papers
- Markov Balance Satisfaction Improves Performance in Strictly Batch Offline Imitation LearningRishabh Agrawal, Nathan Dahlin, Rahul Jain, Ashutosh NayyarAAAI 2025 · 1 citation
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 6 citations
- Imitation with Neural Density ModelsKuno Kim, Akshat Jindal, Yang Song, Jiaming Song et al.NeurIPS 2021 · 14 citations
- Normalizing Flows are Capable Models for Continuous ControlRaj Ghugare, Benjamin EysenbachNeurIPS 2025
- FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot ManipulationQinglun Zhang, Zhen Liu, Haoqiang Fan, Guanghui Liu et al.AAAI 2025 · 5 citations
