Adversarial Bandits Policy for Crawling Commercial Web Content
Shuguang Han, Michael Bendersky, Przemek Gajda, Sergey Novikov, Marc Najork, Bernhard Brodowsky, Alexandrin Popescul
摘要
The rapid growth of commercial web content has driven the development of shopping search services to help users find product offers. Due to the dynamic nature of commercial content, an effective recrawl policy is a key component in a shopping search service; it ensures that users have access to the up-to-date product details. Most of the existing strategies either relied on simple heuristics, or overlooked the resource budgets. To address this, Azar et al. [5] recently proposed an optimization strategy LambdaCrawl aiming to maximize content freshness within a given resource budget. In this paper, we demonstrate that the effectiveness of LambdaCrawl is governed in large part by how well future content change rate can be estimated. By adopting the state-of-the-art deep learning models for change rate prediction, we obtain a substantial increase of content freshness over the common LambdaCrawl implementation with change rate estimated from the past history. Moreover, we demonstrate that while LambdaCrawl is a significant advancement upon existing recrawl strategies, it can be further improved upon by a unified multi-strategy recrawl policy. To this end, we adopt the K-armed adversarial bandits algorithm that can provably optimize the overall freshness by combining multiple strategies. Empirical results over a large-scale production dataset confirm its superiority to LambdaCrawl, especially under tight resource budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- A Scalable Crawling Algorithm Utilizing Noisy Change-Indicating SignalsJulian Zimmert, Róbert Busa-Fekete, András György, Linhai Qiu 等WWW 2025
- RLPer: A Reinforcement Learning Model for Personalized SearchJing Yao, Zhicheng Dou, Jun Xu, Ji-Rong WenWWW 2020 · 被引用 33 次
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou 等ACL 2026 · 被引用 1,216 次
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang 等AAAI 2021 · 被引用 131 次
- Memorize, Factorize, or be Naive: Learning Optimal Feature Interaction Methods for CTR PredictionFuyuan Lyu, Xing Tang, Huifeng Guo, Ruiming Tang 等ICDE 2022 · 被引用 18 次
