AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution Generalization
Xiao Zhang, Sunhao Dai, Jun Xu, Yong Liu, Zhenhua Dong
摘要
Online to batch conversion involves constructing a new batch learner by utilizing a series of models generated by an existing online learning algorithm, for achieving generalization guarantees under i.i.d assumption. However, when applied to real-world streaming applications such as streaming recommender systems, the data stream may be sampled from time-varying distributions instead of persistently being i.i.d. This poses a challenge in terms of out-of-distribution (OOD) generalization. Existing approaches employ fixed conversion mechanisms that are unable to adapt to novel testing distributions, hindering the testing accuracy of the batch learner. To address these issues, we propose AdaO2B, an adaptive online to batch conversion approach under the bandit setting. AdaO2B is designed to be aware of the distribution shifts in the testing data and achieves OOD generalization guarantees. Specifically, AdaO2B can dynamically combine the sequence of models learned by a contextual bandit algorithm and determine appropriate combination weights using a context-aware weighting function. This innovative approach allows for the conversion of a sequence of models into a batch learner that facilitates OOD generalization. Theoretical analysis provides justification for why and how the learned adaptive batch learner can achieve OOD generalization error guarantees. Experimental results have demonstrated that AdaO2B significantly outperforms state-of-the-art baselines on both synthetic and real-world recommendation datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang 等NeurIPS 2021 · 被引用 610 次
- Regret Bounds for Batched BanditsHossein Esfandiari, Amin Karbasi, Abbas Mehrabian, Vahab S. MirrokniAAAI 2021 · 被引用 74 次
- Counterfactual Reward Modification for Streaming Recommendation with Delayed FeedbackXiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang 等SIGIR 2021 · 被引用 60 次
- Impact of Representation Learning in Linear BanditsJiaqi Yang, Wei Hu, Jason D. Lee, Simon Shaolei DuICLR 2021 · 被引用 58 次
- Counteracting User Attention Bias in Music Streaming Recommendation via Reward ModificationXiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong 等KDD 2022 · 被引用 26 次
相关 Paper
- From Online to Non-i.i.d. Batch LearningYufei Tao, Shangqi LuKDD 2020 · 被引用 2 次
- A Generic Learning Framework for Sequential Recommendation with Distribution ShiftsZhengyi Yang, Xiangnan He, Jizhi Zhang, Jiancan Wu 等SIGIR 2023 · 被引用 55 次
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 被引用 241 次
- Test-time Adaptation in Non-stationary Environments via Adaptive Representation AlignmentZhen-Yu Zhang, Zhiyu Xie, Huaxiu Yao, Masashi SugiyamaNeurIPS 2024 · 被引用 12 次
- Reconsidering Learning Objectives in Unbiased Recommendation: A Distribution Shift PerspectiveTeng Xiao, Zhengyu Chen, Suhang WangKDD 2023 · 被引用 8 次
