Imitation Beyond Expectation Using Pluralistic Stochastic Dominance
Ali Farajzadeh, Danyal Saeed, Syed M. Abbas, Rushit N. Shah, Aadirupa Saha, Brian D. Ziebart
摘要
Imitation learning seeks to estimate policies reflecting the values of demonstrated behaviors. Prevalent approaches learn to match or exceed the demonstrator's performance in expectation without knowing the demonstrator's reward function. Unfortunately, this does not induce pluralistic imitators that learn to support distinct demonstrations. We reformulate imitation learning using stochastic dominance over the demonstrations' reward distribution across a range of reward functions as our foundational aim. Our approach matches imitator policy samples (or support) with demonstrations using optimal transport theory to define an imitation learning objective over trajectory pairs. We demonstrate the benefits of pluralistic stochastic dominance (PSD) for imitation in both theory and practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 被引用 90 次
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 被引用 73 次
- Cross-Domain Imitation Learning via Optimal TransportArnaud Fickinger, Samuel Cohen, Stuart Russell, Brandon AmosICLR 2022 · 被引用 65 次
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 被引用 43 次
相关 Paper
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal TransportMingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin WangICML 2025
- Towards Uniformly Superhuman Autonomy via Subdominance MinimizationBrian D. Ziebart, Sanjiban Choudhury, Xinyan Yan, Paul VernazaICML 2022 · 被引用 2 次
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 被引用 1 次
- Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation LearningHanlin Yang, Jian Yao, Weiming Liu, Qing Wang 等ICLR 2025
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 被引用 17 次
