Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback
Michelle D. Zhao, Henny Admoni, Reid G. Simmons, Aaditya Ramdas, Andrea Bajcsy
摘要
In interactive imitation learning (IL), uncertainty quantification offers a way for the learner (i.e. robot) to contend with distribution shifts encountered during deployment by actively seeking additional feedback from an expert (i.e. human) online. Prior works use mechanisms like ensemble disagreement or Monte Carlo dropout to quantify when black-box IL policies are uncertain; however, these approaches can lead to overconfident estimates when faced with deployment-time distribution shifts. Instead, we contend that we need uncertainty quantification algorithms that can leverage the expert human feedback received during deployment time to adapt the robot's uncertainty online. To tackle this, we draw upon online conformal prediction, a distribution-free method for constructing prediction intervals online given a stream of ground-truth labels. Human labels, however, are intermittent in the interactive IL setting. Thus, from the conformal prediction side, we introduce a novel uncertainty quantification algorithm called intermittent quantile tracking (IQT) that leverages a probabilistic model of intermittent labels, maintains asymptotic coverage guarantees, and empirically achieves desired coverage levels. From the interactive IL side, we develop ConformalDAgger, a new approach wherein the robot uses prediction intervals calibrated by IQT as a reliable measure of deployment-time uncertainty to actively query for more expert feedback. We compare ConformalDAgger to prior uncertainty-aware DAgger methods in scenarios where the distribution shift is (and isn't) present because of changes in the expert's policy. We find that in simulated and hardware deployments on a 7DOF robotic manipulator, ConformalDAgger detects high uncertainty when the expert shifts and increases the number of interventions compared to baselines, allowing the robot to more quickly learn the new behavior. Project page at cmu-intentlab.github.io/conformalized-interactive-il/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho 等ICML 2026 · 被引用 3 次
- EAPO: Enhancing Policy Optimization with On-Demand Expert AssistanceSiyao Song, Cong Ma, Zhihao Cheng, Shiye Lei 等ICML 2026
它引用的顶会 Paper11
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Adaptive Conformal Inference Under Distribution ShiftIsaac Gibbs, Emmanuel J. CandèsNeurIPS 2021 · 被引用 665 次
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 被引用 586 次
- Adaptive Conformal Predictions for Time SeriesMargaux Zaffran, Olivier Féron, Yannig Goude, Julie Josse 等ICML 2022 · 被引用 209 次
- Conformal PID Control for Time Series PredictionAnastasios Angelopoulos, Emmanuel J. Candès, Ryan J. TibshiraniNeurIPS 2023 · 被引用 164 次
相关 Paper
- Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret MinimizationJunyoung Yang, Kyungmin Kim, Sangdon ParkICLR 2026 · 被引用 4 次
- Conformal Prediction Beyond the Horizon: Distribution-Free Inference for Policy EvaluationFeichen Gan, Youcun Lu, Yingying Zhang, Yukun LiuNeurIPS 2025 · 被引用 2 次
- Distribution-informed Online Conformal PredictionDongjian Hu, Junxi Wu, Shu-Tao Xia, Changliang ZouICLR 2026 · 被引用 2 次
- Improved Online Conformal Prediction via Strongly Adaptive Online LearningAadyot Bhatnagar, Huan Wang, Caiming Xiong, Yu BaiICML 2023 · 被引用 87 次
- Efficient Online Set-valued Classification with Bandit FeedbackZhou Wang, Xingye QiaoICML 2024 · 被引用 2 次
