BLOB: A Probabilistic Model for Recommendation that Combines Organic and Bandit Signals
Otmane Sakhi, Stephen Bonner, David Rohde, Flavian Vasile
Abstract
A common task for recommender systems is to build a profile of the interests of a user from items in their browsing history and later to recommend items to the user from the same catalog. The users' behavior consists of two parts: the sequence of items that they viewed without intervention (the organic part) and the sequences of items recommended to them and their outcome (the bandit part). In this paper, we propose Bayesian Latent Organic Bandit model (BLOB), a probabilistic approach to combine the 'organic' and 'bandit' signals in order to improve the estimation of recommendation quality. The bandit signal is valuable as it gives direct feedback of recommendation performance, but the signal quality is very uneven, as it is highly concentrated on the recommendations deemed optimal by the past version of the recommender system. In contrast, the organic signal is typically strong and covers most items, but is not always relevant to the recommendation task. In order to leverage the organic signal to efficiently learn the bandit signal in a Bayesian model we identify three fundamental types of distances, namely action-history, action-action and history-history distances. We implement a scalable approximation of the full model using variational auto-encoders and the local re-paramerization trick. We show using extensive simulation studies that our method out-performs or matches the value of both state-of-the-art organic-based recommendation algorithms, and of bandit-based methods (both value and policy-based) both in organic and bandit-rich environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51d46a94-b4d0-44bd-b93a-11e6fd5a6882Cited by top-tier papers9
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 24 citations
- PAC-Bayesian Offline Contextual Bandits With GuaranteesOtmane Sakhi, Pierre Alquier, Nicolas ChopinICML 2023 · 23 citations
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Exponential Smoothing for Off-Policy LearningImad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna KorbaICML 2023 · 17 citations
- On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-n RecommendationOlivier Jeunen, Ivan Potapov, Aleksei UstimenkoKDD 2024 · 16 citations
Builds on1
Related papers
- An Epistemic Position-Based Click Model: From Interactions to Epistemic Distributions of Relevance and BiasOscar Rolando Ramirez Milian, Harrie OosterhuisSIGIR 2026 · 1 citation
- Unifying Behavior Modeling and Semantic Generation for Generative RecommendationBinquan Wu, Xinbo Chen, Yicheng Luo, Yuhao Ke et al.KDD 2026
- Impatient Bandits: Optimizing Recommendations for the Long-Term Without DelayThomas M. McDonald, Lucas Maystre, Mounia Lalmas, Daniel Russo et al.KDD 2023 · 12 citations
- EAGER: Two-Stream Generative Recommender with Behavior-Semantic CollaborationYe Wang, Jiahao Xun, Minjie Hong, Jieming Zhu et al.KDD 2024 · 13 citations
- M²VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item RecommendationChuan He, Yongchao Liu, Qiang Li, Chuntao Hong et al.AAAI 2026 · 1 citation
