Context-Aware Document Term Weighting for Ad-Hoc Search
Zhuyun Dai, Jamie Callan
Abstract
Bag-of-words document representations play a fundamental role in modern search engines, but their power is limited by the shallow frequency-based term weighting scheme. This paper proposes HDCT, a context-aware document term weighting framework for document indexing and retrieval. It first estimates the semantic importance of a term in the context of each passage. These fine-grained term weights are then aggregated into a document-level bag-of-words representation, which can be stored into a standard inverted index for efficient retrieval. This paper also proposes two approaches that enable training HDCT without relevance labels. Experiments show that an index using HDCT weights significantly improved the retrieval accuracy compared to typical term-frequency and state-of-the-art embedding-based indexes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b209d75-912f-46b9-a42b-cce40205f88eCited by top-tier papers9
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang et al.NeurIPS 2023 · 151 citations
- Efficient Inverted Indexes for Approximate Retrieval over Learned Sparse RepresentationsSebastian Bruch, Franco Maria Nardini, Cosimo Rulli, Rossano VenturiniSIGIR 2024 · 47 citations
- Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span PredictionXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan et al.SIGIR 2022 · 33 citations
- Efficient Neural Ranking using Forward IndexesJurek Leonhardt, Koustav Rudra, Megha Khosla, Abhijit Anand et al.WWW 2022 · 16 citations
Related papers
- Compact Token Representations with Contextual Quantization for Efficient Document Re-rankingYingrui Yang, Yifan Qiao, Tao YangACL 2022 · 8 citations
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document EmbeddingsMax Conti, Manuel Faysse, Gautier Viaud, Antoine Bosselut et al.EMNLP 2025 · 1 citation
- LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale RetrievalTao Shen, Xiubo Geng, Chongyang Tao, Can Xu et al.ICLR 2023 · 14 citations
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 5 citations
- Query-as-context Pre-training for Dense Passage RetrievalXing Wu, Guangyuan Ma, Wanhui Qian, Zijia Lin et al.EMNLP 2023 · 2 citations
