Lune

NeurIPS2025Top-tier venue

Imitation Beyond Expectation Using Pluralistic Stochastic Dominance

Ali Farajzadeh, Danyal Saeed, Syed M. Abbas, Rushit N. Shah, Aadirupa Saha, Brian D. Ziebart

2025Year
2Citations

Abstract

Imitation learning seeks to estimate policies reflecting the values of demonstrated behaviors. Prevalent approaches learn to match or exceed the demonstrator's performance in expectation without knowing the demonstrator's reward function. Unfortunately, this does not induce pluralistic imitators that learn to support distinct demonstrations. We reformulate imitation learning using stochastic dominance over the demonstrations' reward distribution across a range of reward functions as our foundational aim. Our approach matches imitator policy samples (or support) with demonstrations using optimal transport theory to define an imitation learning objective over trajectory pairs. We demonstrate the benefits of pluralistic stochastic dominance (PSD) for imitation in both theory and practice.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cd764cf7-f2be-4238-a2cc-2e342effe4ad

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines