Lune

ICML2026顶会

Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal Dataset

Quang Anh PHAM, Tien Mai, Akshat Kumar

出版方
2026年份

摘要

Imitation Learning (IL) learns high-quality policies from expert demonstrations but degrades in low-expert-data regimes. To address this, recent work studies `` offline IL with supplementary data ", augmenting expert data with trajectories from suboptimal policies. A prominent framework is Distribution Correction Estimation (DICE), which estimates density ratios via the dual of a divergence minimization problem between learned and expert visitation distributions. However, existing DICE-based methods either rely on a strict coverage assumption or introduce additional dataset regularization, limiting performance. We propose ReDICE , a new DICE-based method that addresses these issues through an objective-level reformulation. Our approach constructs a mixture-distribution objective that preserves the original expert-imitation objective while removing the coverage assumption, and its dual reduces to a stable Gumbel regression objective for efficient optimization. We further introduce a novel policy extraction mechanism that improves performance. Experiments on standard and real-world offline IL benchmarks show that ReDICE consistently outperforms prior methods and achieves state-of-the-art results.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖