DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
Rahul Mehta, Kavin R. V, Indrajit Pal, Tushar Abhishek, Pawan Goyal, Manish Gupta
Abstract
Query auto-completion (QAC) has been widely studied in the context of web search, yet remains underexplored for in-document search, which we term DocQAC. DocQAC aims to enhance search productivity within long documents by helping users craft faster, more precise queries, even for complex or hard-to-spell terms. Unlike traditional WebQAC systems, DocQAC can leverage rich document context, having access not only to the partially typed user query and global historical queries, but also the content of the current document itself, and crucially, the document-specific history of user query interactions. To address this setting, we propose a novel adaptive trie-guided decoding framework that uses user query prefixes to softly steer language models toward high-quality completions. Our approach introduces an adaptive penalty mechanism with tunable hyperparameters, enabling a principled trade-off between model confidence and trie-based guidance. To efficiently incorporate document context, we explore retrieval-augmented generation (RAG) and lightweight contextual document signals such as titles, keyphrases, and summaries. When applied to encoder–decoder models like T5 and BART, our trie-guided framework outperforms strong baselines and even surpasses much larger instruction-tuned models such as LLaMA-3 and Phi-3 in seen-query settings. This demonstrates its practicality for real-world DocQAC system deployments, where efficiency and scalability are critical. We evaluate our method on a newly introduced DocQAC benchmark derived from ORCAS, enriched with query–document pairs. We make both the DocQAC dataset https://bit.ly/3IGEkbH and code https://github.com/rahcode7/DocQAC publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04a5ccff-de2e-43ff-8d14-7aa94d3d51b4Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 200 citations
- Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningSaibo Geng, Martin Josifoski, Maxime Peyrard, Robert WestEMNLP 2023 · 33 citations
Related papers
- Grounding Language Model with Chunking-Free In-Context RetrievalHongjin Qian, Zheng Liu, Kelong Mao, Yujia Zhou et al.ACL 2024
- PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual AdapterHaoyan Yang, Zhitao Li, Yong Zhang, Jianzong Wang et al.EMNLP 2023 · 16 citations
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval AugmentationHongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao et al.WWW 2025 · 92 citations
- SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAGXuechen Zhang, Koustava Goswami, Samet Oymak, Jiasi Chen et al.ICLR 2026 · 1 citation
- DVD: Dynamic Contrastive Decoding for Knowledge Amplification in Multi-Document Question AnsweringJing Jin, Houfeng Wang, Hao Zhang, Xiaoguang Li et al.EMNLP 2024 · 1 citation
