Lune

SIGIR2026Top-tier venue

Recasting Web-Scale Query Suggestion as dense retrieval: Efficient, Up-to-Date, and Context-Aware Suggestions

Sosuke Nishikawa, Naoki Yoshinaga, Nobuhiro Kaji

2026Year

Abstract

Query suggestion (QS) in web search must provide efficient and up-to-date suggestions that are relevant to the session context. Recent studies treat QS as generation, which captures session context well but suffers from high latency and costly retraining to maintain information freshness. In this study, we recast QS as dense retrieval: given a search session, the next query is retrieved from a large index of historical queries using efficient approximate nearest neighbor search. We propose Context-aware Asymmetric Dual Encoder for QS (CADE-QS), which uses an asymmetric dual encoder to model query -to- suggestion dependency and incorporates session history into the query encoder via context-aware contrastive learning. We evaluate our method on two real-world search-log datasets: a recent Japanese web search log and the public AOL log. CADE-QS rivals strong generative baselines in quality while reducing end-to-end latency to around 30 ms on CPU, representing an order-of-magnitude improvement. Detailed analyses confirm CADE-QS's robust context awareness, unidirectional modeling, and practicality for reflecting evolving information via index refreshes, as well as the effectiveness of an adaptive hybrid strategy for low-coverage scenarios. Our code is publicly available at https://github.com/lycorp-jp/cadeqs.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 1b3f7bd3-850a-416a-abde-3f4397bdce24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines