Match Plan Generation in Web Search with Parameterized Action Reinforcement Learning
Ziyan Luo, Linfeng Zhao, Wei Cheng, Sihao Chen, Qi Chen, Hui Xue, Haidong Wang, Chuanjie Liu, Mao Yang, Lintao Zhang
Abstract
To achieve good result quality and short query response time, search engines use specific match plans on Inverted Index to help retrieve a small set of relevant documents from billions of web pages. A match plan is composed of a sequence of match rules, which contain discrete match rule types and continuous stopping quotas. Currently, match plans are manually designed by experts according to their several years' experience, which encounters difficulty in dealing with heterogeneous queries and varying data distribution. In this work, we formulate the match plan generation as a Partially Observable Markov Decision Process (POMDP) with a parameterized action space, and propose a novel reinforcement learning algorithm Parameterized Action Soft Actor-Critic (PASAC) to effectively enhance the exploration in both spaces. In our scene, we also discover a skew prioritizing issue of the original Prioritized Experience Replay (PER) and introduce Stratified Prioritized Experience Replay (SPER) to address it. We are the first group to generalize this task for all queries as a learning problem with zero prior knowledge and successfully apply deep reinforcement learning in the real web search environment. Our approach greatly outperforms the welldesigned production match plans by over 70% reduction of index block accesses with the quality of documents almost unchanged, and 9% reduction of query response time even with model inference cost. Our method also beats the baselines on some open-source benchmarks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acd47708-cd0f-4fc0-91b9-cd7110fe8ae9Related papers
- RLPer: A Reinforcement Learning Model for Personalized SearchJing Yao, Zhicheng Dou, Jun Xu, Ji-Rong WenWWW 2020 · 33 citations
- FOSS: A Self-Learned Doctor for Query OptimizerKai Zhong, Luming Sun, Tao Ji, Cuiping Li et al.ICDE 2024 · 5 citations
- Automatic Web Testing Using Curiosity-Driven Reinforcement LearningYan Zheng, Yi Liu, Xiaofei Xie, Yepang Liu et al.ICSE 2021 · 75 citations
- Balsa: Learning a Query Optimizer Without Expert DemonstrationsZongheng Yang, Wei-Lin Chiang, Sifei Luan, Gautam Mittal et al.SIGMOD 2022 · 99 citations
- DBA bandits: Self-driving index tuning under ad-hoc, analytical workloads with safety guaranteesR. Malinga Perera, Bastian Oetomo, Benjamin I. P. Rubinstein, Renata Borovica-GajicICDE 2021 · 40 citations
