Towards Domain Adaptive Neural Contextual Bandits
Ziyan Wang, Xiaoming Huo, Hao Wang
摘要
Contextual bandit algorithms are essential for solving real-world decision making problems. In practice, collecting a contextual bandit's feedback from different domains may involve different costs. For example, measuring drug reaction from mice (as a source domain) and humans (as a target domain). Unfortunately, adapting a contextual bandit algorithm from a source domain to a target domain with distribution shift still remains a major challenge and largely unexplored. In this paper, we introduce the first general domain adaptation method for contextual bandits. Our approach learns a bandit model for the target domain by collecting feedback from the source domain. Our theoretical analysis shows that our algorithm maintains a sub-linear regret bound even adapting across domains. Empirical results show that our approach outperforms the state-of-the-art contextual bandit algorithms on real-world datasets. Code will soon be available at https:// github.com/Wang-ML-Lab/DABand .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- SENTRY: Selective Entropy Optimization via Committee Consistency for Unsupervised Domain AdaptationViraj Prabhu, Shivam Khare, Deeksha Kartik, Judy HoffmanICCV 2021 · 被引用 155 次
- Continuously Indexed Domain AdaptationHao Wang, Hao He, Dina KatabiICML 2020 · 被引用 129 次
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-BanditsHan Guo, Ramakanth Pasunuru, Mohit BansalAAAI 2020 · 被引用 120 次
- Neural Contextual Bandits with Deep Representation and Shallow ExplorationPan Xu, Zheng Wen, Handong Zhao, Quanquan GuICLR 2022 · 被引用 90 次
相关 Paper
- Adversarial Attacks on Linear Contextual BanditsEvrard Garcelon, Baptiste Rozière, Laurent Meunier, Jean Tarbouriech 等NeurIPS 2020 · 被引用 60 次
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 被引用 241 次
- Neural Dueling Bandits: Preference-Based Optimization with Human FeedbackArun Verma, Zhongxiang Dai, Xiaoqiang Lin, Patrick Jaillet 等ICLR 2025
- Model Transferability with Responsive Decision SubjectsYatong Chen, Zeyu Tang, Kun Zhang, Yang LiuICML 2023 · 被引用 11 次
- Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsTong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma BrunskillNeurIPS 2022 · 被引用 8 次
