AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution Generalization
Xiao Zhang, Sunhao Dai, Jun Xu, Yong Liu, Zhenhua Dong
Abstract
Online to batch conversion involves constructing a new batch learner by utilizing a series of models generated by an existing online learning algorithm, for achieving generalization guarantees under i.i.d assumption. However, when applied to real-world streaming applications such as streaming recommender systems, the data stream may be sampled from time-varying distributions instead of persistently being i.i.d. This poses a challenge in terms of out-of-distribution (OOD) generalization. Existing approaches employ fixed conversion mechanisms that are unable to adapt to novel testing distributions, hindering the testing accuracy of the batch learner. To address these issues, we propose AdaO2B, an adaptive online to batch conversion approach under the bandit setting. AdaO2B is designed to be aware of the distribution shifts in the testing data and achieves OOD generalization guarantees. Specifically, AdaO2B can dynamically combine the sequence of models learned by a contextual bandit algorithm and determine appropriate combination weights using a context-aware weighting function. This innovative approach allows for the conversion of a sequence of models into a batch learner that facilitates OOD generalization. Theoretical analysis provides justification for why and how the learned adaptive batch learner can achieve OOD generalization error guarantees. Experimental results have demonstrated that AdaO2B significantly outperforms state-of-the-art baselines on both synthetic and real-world recommendation datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d48e2095-efa6-419c-b696-18d929b2d6c9Cited by top-tier papers1
Ask how each one uses itBuilds on7
- R-Drop: Regularized Dropout for Neural NetworksXiaobo Liang, Lijun Wu, Juntao Li, Yue Wang et al.NeurIPS 2021 · 610 citations
- Regret Bounds for Batched BanditsHossein Esfandiari, Amin Karbasi, Abbas Mehrabian, Vahab S. MirrokniAAAI 2021 · 74 citations
- Counterfactual Reward Modification for Streaming Recommendation with Delayed FeedbackXiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang et al.SIGIR 2021 · 60 citations
- Impact of Representation Learning in Linear BanditsJiaqi Yang, Wei Hu, Jason D. Lee, Simon Shaolei DuICLR 2021 · 58 citations
- Counteracting User Attention Bias in Music Streaming Recommendation via Reward ModificationXiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong et al.KDD 2022 · 26 citations
Related papers
- From Online to Non-i.i.d. Batch LearningYufei Tao, Shangqi LuKDD 2020 · 2 citations
- A Generic Learning Framework for Sequential Recommendation with Distribution ShiftsZhengyi Yang, Xiangnan He, Jizhi Zhang, Jiancan Wu et al.SIGIR 2023 · 55 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
- Test-time Adaptation in Non-stationary Environments via Adaptive Representation AlignmentZhen-Yu Zhang, Zhiyu Xie, Huaxiu Yao, Masashi SugiyamaNeurIPS 2024 · 12 citations
- Reconsidering Learning Objectives in Unbiased Recommendation: A Distribution Shift PerspectiveTeng Xiao, Zhengyu Chen, Suhang WangKDD 2023 · 8 citations
