What Are You Token About? Dense Retrieval as Distributions Over the Vocabulary
Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson
Abstract
Dual encoders are now the dominant architecture for dense retrieval. Yet, we have little understanding of how they represent text, and why this leads to good performance. In this work, we shed light on this question via distributions over the vocabulary. We propose to interpret the vector representations produced by dual encoders by projecting them into the model's vocabulary space. We show that the resulting projections contain rich semantic information, and draw connection between them and sparse retrieval. We find that this view can offer an explanation for some of the failure cases of dense retrievers. For example, we observe that the inability of models to handle tail entities is correlated with a tendency of the token distributions to forget some of the tokens of those entities. We leverage this insight and propose a simple way to enrich query and passage representations with lexical information at inference time, and show that this significantly improves performance compared to the original model in zero-shot settings, and specifically on the BEIR benchmark. 1 * Supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion. 1 Our code is publicly available at https://github. com/oriram/dense-retrieval-projections . Uninterpretable Q: Where was Michael Jack born? Query Encoder MLM Head michael -0.54 jack -1.39 son -4.19 father -4.58 birth -4.83 boy -5.15 family -5.75 locke -5.78 childhood -6.10 baby -6.11 child -6.15 ⋮ Interpretable CLS where was michael jack born ? 1. michael -0.54 2. jack -1.39 3. son -4.19 4. father -4.58 5. birth -4.83 6. boy -5.15 7. family -5.75 ⋮ Michael Jack, (born 17 September 1946) is a Conservative Party politician in the United Kingdom … Michael Jack was born in Folkestone, Kent, England. Q: How many judges currently serve on the Supreme Court? Query Encoder MLM Head 1.court -1.43 2. judges -1.71 3. justices -2.27 4. judge -2.96 5. judicial -3.52 6. nine -3.81 7. courts -4.33 ⋮ Passage Encoder MLM Head Demographics of the Supreme Court of the United States … In 2008, seven of the nine sitting justices were millionaires 1. he -2.42 2. jack -3.48 3. major -3.83 4. labour -3.84 5. chairman -3.92 ⋮ 146. michael -6.97 ⋮ 1. justices -0.71 2. court -2.82 3. judges -3.11 4. judge -3.74 5. judicial -4.16 ⋮ 20.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0862dabf-d4ba-486f-86c0-b8bd7fc4b6f8Cited by top-tier papers15
- In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination MitigationShiqi Chen, Miao Xiong, Junteng Liu, Zhengxuan Wu et al.ICML 2024 · 49 citations
- The Hidden Language of Diffusion ModelsHila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin et al.ICLR 2024 · 38 citations
- Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsMor Geva, Jasmijn Bastings, Katja Filippova, Amir GlobersonEMNLP 2023 · 38 citations
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 5 citations
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 5 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
Related papers
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- LED: Lexicon-Enlightened Dense Retriever for Large-Scale RetrievalKai Zhang, Chongyang Tao, Tao Shen, Can Xu et al.WWW 2023 · 27 citations
- Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense RetrievalChaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao et al.ACL 2024 · 10 citations
- BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense RetrievalShicheng Xu, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 8 citations
