Factored DRO: Factored Distributionally Robust Policies for Contextual Bandits
Tong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma Brunskill
摘要
While there has been extensive work on learning from offline data for contextual multi-armed bandit settings, existing methods typically assume there is no environment shift: that the learned policy will operate in the same environmental process as that of data collection. However, this assumption may limit the use of these methods for many practical situations where there may be distribution shifts. In this work we propose Factored Distributionally Robust Optimization (Factored-DRO) 1 , which is able to separately handle distribution shifts in the context distribution and shifts in the reward generating process. Prior work that either ignores potential shifts in the context, or considers them jointly, can lead to performance that is too conservative, especially under certain forms of reward feedback. Our Factored-DRO objective mitigates this by considering the shifts separately, and our proposed estimators are consistent and converge asymptotically. We also introduce a practical algorithm and demonstrate promising empirical results in environments based on real-world datasets, such as voting outcomes and scene classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Not all distributional shifts are equal: Fine-grained robust conformal inferenceJiahao Ai, Zhimei RenICML 2024 · 被引用 15 次
- Distributionally Robust Optimization with Bias and Variance ReductionRonak Mehta, Vincent Roulet, Krishna Pillutla, Zaïd HarchaouiICLR 2024 · 被引用 6 次
- Distributionally Robust Policy Learning under Concept DriftsJingyuan Wang, Zhimei Ren, Ruohan Zhan, Zhengyuan ZhouICML 2025
- Off-Policy Evaluation and Learning for the Future under Non-StationarityTatsuhiro Shimizu, Kazuki Kawamura, Takanori Muroi, Yusuke Narita 等KDD 2025
它引用的顶会 Paper10
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 125 次
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White 等ICML 2020 · 被引用 72 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 被引用 60 次
相关 Paper
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 被引用 1 次
- Distributionally Robust Policy Evaluation and Learning in Offline Contextual BanditsNian Si, Fan Zhang, Zhengyuan Zhou, Jose H. BlanchetICML 2020 · 被引用 59 次
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova 等AAAI 2020 · 被引用 48 次
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 被引用 35 次
- Doubly Robust Distributionally Robust Offline Contextual PricingMin Xu, Xinyi Yin, Yunfan Zhang, Yuxuan Han 等ICML 2026 · 被引用 15 次
