Off-Policy Evaluation for Human Feedback
Qitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh, Min Chi, Miroslav Pajic
摘要
Off-policy evaluation (OPE) is important for closing the gap between offline training and evaluation of reinforcement learning (RL), by estimating performance and/or rank of target (evaluation) policies using offline trajectories only. It can improve the safety and efficiency of data collection and policy testing procedures in situations where online deployments are expensive, such as healthcare. However, existing OPE methods fall short in estimating human feedback (HF) signals, as HF may be conditioned over multiple underlying factors and is only sparsely available; as opposed to the agent-defined environmental rewards (used in policy optimization), which are usually determined over parametric functions or distributions. Consequently, the nature of HF signals makes extrapolating accurate OPE estimations to be challenging. To resolve this, we introduce an OPE for HF (OPEHF) framework that revives existing OPE methods in order to accurately evaluate the HF signals. Specifically, we develop an immediate human reward (IHR) reconstruction approach, regularized by environmental knowledge distilled in a latent space that captures the underlying dynamics of state transitions as well as issuing HF signals. Our approach has been tested over two real-world experiments, adaptive in-vivo neurostimulation and intelligent tutoring, as well as in a simulation environment (visual Q&A). Results show that our approach significantly improves the performance toward estimating HF signals accurately, compared to directly applying (variants of) existing OPE methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju 等ICLR 2024 · 被引用 5 次
- Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent TutoringGe Gao, Xi Yang, Min ChiAAAI 2024 · 被引用 1 次
- OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language ModelsJunda Wu, Xintong Li, Ruoyu Wang, Yu Xia 等ICLR 2025
它引用的顶会 Paper23
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
相关 Paper
- Counterfactual-Augmented Importance Sampling for Semi-Offline Policy EvaluationShengpu Tang, Jenna WiensNeurIPS 2023 · 被引用 8 次
- A Maximum-Entropy Approach to Off-Policy Evaluation in Average-Reward MDPsNevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine 等NeurIPS 2020 · 被引用 13 次
- Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data UtilizationYihan Du, Anna Winnicki, Gal Dalal, Shie Mannor 等ICML 2024 · 被引用 22 次
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
- Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human FeedbackYifu Yuan, Jianye Hao, Yi Ma, Zibin Dong 等ICLR 2024 · 被引用 21 次
