Match Plan Generation in Web Search with Parameterized Action Reinforcement Learning
Ziyan Luo, Linfeng Zhao, Wei Cheng, Sihao Chen, Qi Chen, Hui Xue, Haidong Wang, Chuanjie Liu, Mao Yang, Lintao Zhang
摘要
To achieve good result quality and short query response time, search engines use specific match plans on Inverted Index to help retrieve a small set of relevant documents from billions of web pages. A match plan is composed of a sequence of match rules, which contain discrete match rule types and continuous stopping quotas. Currently, match plans are manually designed by experts according to their several years' experience, which encounters difficulty in dealing with heterogeneous queries and varying data distribution. In this work, we formulate the match plan generation as a Partially Observable Markov Decision Process (POMDP) with a parameterized action space, and propose a novel reinforcement learning algorithm Parameterized Action Soft Actor-Critic (PASAC) to effectively enhance the exploration in both spaces. In our scene, we also discover a skew prioritizing issue of the original Prioritized Experience Replay (PER) and introduce Stratified Prioritized Experience Replay (SPER) to address it. We are the first group to generalize this task for all queries as a learning problem with zero prior knowledge and successfully apply deep reinforcement learning in the real web search environment. Our approach greatly outperforms the welldesigned production match plans by over 70% reduction of index block accesses with the quality of documents almost unchanged, and 9% reduction of query response time even with model inference cost. Our method also beats the baselines on some open-source benchmarks 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- RLPer: A Reinforcement Learning Model for Personalized SearchJing Yao, Zhicheng Dou, Jun Xu, Ji-Rong WenWWW 2020 · 被引用 33 次
- FOSS: A Self-Learned Doctor for Query OptimizerKai Zhong, Luming Sun, Tao Ji, Cuiping Li 等ICDE 2024 · 被引用 5 次
- Automatic Web Testing Using Curiosity-Driven Reinforcement LearningYan Zheng, Yi Liu, Xiaofei Xie, Yepang Liu 等ICSE 2021 · 被引用 75 次
- Balsa: Learning a Query Optimizer Without Expert DemonstrationsZongheng Yang, Wei-Lin Chiang, Sifei Luan, Gautam Mittal 等SIGMOD 2022 · 被引用 99 次
- DBA bandits: Self-driving index tuning under ad-hoc, analytical workloads with safety guaranteesR. Malinga Perera, Bastian Oetomo, Benjamin I. P. Rubinstein, Renata Borovica-GajicICDE 2021 · 被引用 40 次
