Distributionally Robust Policy Evaluation and Learning in Offline Contextual Bandits
Nian Si, Fan Zhang, Zhengyuan Zhou, Jose H. Blanchet
摘要
Policy learning using historical observational data is an important problem that has found widespread applications. However, existing literature rests on the crucial assumption that the future environment where the learned policy will be deployed is the same as the past environment that has generated the data-an assumption that is often false or too coarse an approximation. In this paper, we lift this assumption and aim to learn a distributionally robust policy with bandit observational data. We propose a novel learning algorithm that is able to learn a robust policy to adversarial perturbations and unknown covariate shifts. We first present a policy evaluation procedure in the ambiguous environment and also give a heuristic algorithm to solve the distributionally robust policy learning problems efficiently. Additionally, we provide extensive simulations to demonstrate the robustness of our policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 被引用 104 次
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong 等ICML 2022 · 被引用 72 次
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 被引用 39 次
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil 等NeurIPS 2025 · 被引用 18 次
- Robust Bandit Learning with Imperfect ContextJianyi Yang, Shaolei RenAAAI 2021 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 被引用 1 次
- Distributionally Robust Policy Learning under Concept DriftsJingyuan Wang, Zhimei Ren, Ruohan Zhan, Zhengyuan ZhouICML 2025
- Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsTong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma BrunskillNeurIPS 2022 · 被引用 8 次
- Learning Robust Decision Policies from Observational DataMuhammad Osama, Dave Zachariah, Peter StoicaNeurIPS 2020 · 被引用 6 次
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 被引用 10 次
