Effective and Efficient Query-aware Snippet Extraction for Web Search
Jingwei Yi, Fangzhao Wu, Chuhan Wu, Xiaolong Huang, Binxing Jiao, Guangzhong Sun, Xing Xie
Abstract
Query-aware webpage snippet extraction is widely used in search engines to help users better understand the content of the returned webpages before clicking. The extracted snippet is expected to summarize the webpage in the context of the input query. Existing snippet extraction methods mainly rely on handcrafted features of overlapping words, which cannot capture deep semantic relationships between the query and webpages. Another idea is to extract the sentences which are most relevant to queries as snippets with existing text matching methods. However, these methods ignore the contextual information of webpages, which may be sub-optimal. In this paper, we propose an effective query-aware webpage snippet extraction method named DeepQSE. In DeepQSE, the concatenation of title, query and each candidate sentence serves as an input of query-aware sentence encoder, aiming to capture the fine-grained relevance between the query and sentences. Then, these query-aware sentence representations are modeled jointly through a document-aware relevance encoder to capture contextual information of the webpage. Since the query and each sentence are jointly modeled in DeepQSE, its online inference may be slow. Thus, we further propose an efficient version of DeepQSE, named Efficient-DeepQSE, which can significantly improve the inference speed of DeepQSE without affecting its performance. The core idea of Efficient-DeepQSE is to decompose the query-aware snippet extraction task into two stages, i.e., a coarse-grained candidate sentence selection stage where sentence representations can be cached, and a fine-grained relevance modeling stage. Experiments on two datasets validate the effectiveness and efficiency of our methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4f55a1a-78a9-4e4a-b38e-700dfba91e40Builds on3
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Abstractive Snippet GenerationWei-Fan Chen, Shahbaz Syed, Benno Stein, Matthias Hagen et al.WWW 2020 · 33 citations
Related papers
- Preserve Context Information for Extract-Generate Long-Input Summarization FrameworkRuifeng Yuan, Zili Wang, Ziqiang Cao, Wenjie LiAAAI 2023 · 3 citations
- A Graph-based Relevance Matching Model for Ad-hoc RetrievalYufeng Zhang, Jinghao Zhang, Zeyu Cui, Shu Wu et al.AAAI 2021 · 26 citations
- QaVA: Query-Aware Video Analysis Framework Based on Data Access PatternTianxiong Zhong, Zhiwei Zhang, Yihang Fu, Guo Lu et al.ICDE 2025 · 1 citation
- Personalized News Recommendation with Knowledge-aware Interactive MatchingTao Qi, Fangzhao Wu, Chuhan Wu, Yongfeng HuangSIGIR 2021 · 82 citations
- BERT-ER: Query-specific BERT Entity Representations for Entity RankingShubham Chatterjee, Laura DietzSIGIR 2022 · 18 citations
