Imitation Beyond Expectation Using Pluralistic Stochastic Dominance
Ali Farajzadeh, Danyal Saeed, Syed M. Abbas, Rushit N. Shah, Aadirupa Saha, Brian D. Ziebart
Abstract
Imitation learning seeks to estimate policies reflecting the values of demonstrated behaviors. Prevalent approaches learn to match or exceed the demonstrator's performance in expectation without knowing the demonstrator's reward function. Unfortunately, this does not induce pluralistic imitators that learn to support distinct demonstrations. We reformulate imitation learning using stochastic dominance over the demonstrations' reward distribution across a range of reward functions as our foundational aim. Our approach matches imitator policy samples (or support) with demonstrations using optimal transport theory to define an imitation learning objective over trajectory pairs. We demonstrate the benefits of pluralistic stochastic dominance (PSD) for imitation in both theory and practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd764cf7-f2be-4238-a2cc-2e342effe4adBuilds on10
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 73 citations
- Cross-Domain Imitation Learning via Optimal TransportArnaud Fickinger, Samuel Cohen, Stuart Russell, Brandon AmosICLR 2022 · 65 citations
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 43 citations
Related papers
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal TransportMingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin WangICML 2025
- Towards Uniformly Superhuman Autonomy via Subdominance MinimizationBrian D. Ziebart, Sanjiban Choudhury, Xinyan Yan, Paul VernazaICML 2022 · 2 citations
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 1 citation
- Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation LearningHanlin Yang, Jian Yao, Weiming Liu, Qing Wang et al.ICLR 2025
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 17 citations
