Conformal Prediction Beyond the Horizon: Distribution-Free Inference for Policy Evaluation
Feichen Gan, Youcun Lu, Yingying Zhang, Yukun Liu
摘要
Reliable uncertainty quantification is crucial for reinforcement learning (RL) in high-stakes settings. We propose a unified conformal prediction framework for infinite-horizon policy evaluation that constructs distribution-free prediction intervals for returns in both on-policy and off-policy settings. Our method integrates distributional RL with conformal calibration, addressing challenges such as unobserved returns, temporal dependencies, and distributional shifts. We propose a modular pseudo-return construction based on truncated rollouts and a time-aware calibration strategy using experience replay and weighted subsampling. These innovations mitigate model bias and restore approximate exchangeability, enabling uncertainty quantification even under policy shifts. Our theoretical analysis provides coverage guarantees that account for model misspecification and importance weight estimation. Empirical results, including experiments in synthetic and benchmark environments like Mountain Car, show that our method significantly improves coverage and reliability over standard distributional RL baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 53 次
- Conformal Predictions under Markovian DataFrédéric Zheng, Alexandre ProutièreICML 2024 · 被引用 3 次
- Wasserstein-Regularized Conformal Prediction under General Distribution ShiftRui Xu, Chao Chen, Yue Sun, Parvathinathan Venkitasubramaniam 等ICLR 2025
- Risk-Aware Reinforcement Learning with Coherent Risk Measures and Non-linear Function ApproximationThanh Lam, Arun Verma, Bryan Kian Hsiang Low, Patrick JailletICLR 2023
相关 Paper
- Non-exchangeable Conformal Prediction with Optimal Transport: Tackling Distribution Shift with Unlabeled DataAlvaro H. C. Correia, Christos LouizosNeurIPS 2025 · 被引用 5 次
- Model Uncertainty Quantification by Conformal Prediction in Continual LearningRui Gao, Weiwei LiuICML 2025
- Conformal Time-series ForecastingKamile Stankeviciute, Ahmed M. Alaa, Mihaela van der SchaarNeurIPS 2021 · 被引用 233 次
- Domain-Shift-Aware Conformal Prediction for Large Language ModelsZhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao 等ICML 2026 · 被引用 6 次
- Adaptive Conformal Prediction Intervals for Invariant LearningShuxin Liang, Yihan Xiao, Linglong Kong, Wenlu TangKDD 2025
