What Are You Token About? Dense Retrieval as Distributions Over the Vocabulary
Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson
摘要
Dual encoders are now the dominant architecture for dense retrieval. Yet, we have little understanding of how they represent text, and why this leads to good performance. In this work, we shed light on this question via distributions over the vocabulary. We propose to interpret the vector representations produced by dual encoders by projecting them into the model's vocabulary space. We show that the resulting projections contain rich semantic information, and draw connection between them and sparse retrieval. We find that this view can offer an explanation for some of the failure cases of dense retrievers. For example, we observe that the inability of models to handle tail entities is correlated with a tendency of the token distributions to forget some of the tokens of those entities. We leverage this insight and propose a simple way to enrich query and passage representations with lexical information at inference time, and show that this significantly improves performance compared to the original model in zero-shot settings, and specifically on the BEIR benchmark. 1 * Supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion. 1 Our code is publicly available at https://github. com/oriram/dense-retrieval-projections . Uninterpretable Q: Where was Michael Jack born? Query Encoder MLM Head michael -0.54 jack -1.39 son -4.19 father -4.58 birth -4.83 boy -5.15 family -5.75 locke -5.78 childhood -6.10 baby -6.11 child -6.15 ⋮ Interpretable CLS where was michael jack born ? 1. michael -0.54 2. jack -1.39 3. son -4.19 4. father -4.58 5. birth -4.83 6. boy -5.15 7. family -5.75 ⋮ Michael Jack, (born 17 September 1946) is a Conservative Party politician in the United Kingdom … Michael Jack was born in Folkestone, Kent, England. Q: How many judges currently serve on the Supreme Court? Query Encoder MLM Head 1.court -1.43 2. judges -1.71 3. justices -2.27 4. judge -2.96 5. judicial -3.52 6. nine -3.81 7. courts -4.33 ⋮ Passage Encoder MLM Head Demographics of the Supreme Court of the United States … In 2008, seven of the nine sitting justices were millionaires 1. he -2.42 2. jack -3.48 3. major -3.83 4. labour -3.84 5. chairman -3.92 ⋮ 146. michael -6.97 ⋮ 1. justices -0.71 2. court -2.82 3. judges -3.11 4. judge -3.74 5. judicial -4.16 ⋮ 20.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination MitigationShiqi Chen, Miao Xiong, Junteng Liu, Zhengxuan Wu 等ICML 2024 · 被引用 49 次
- The Hidden Language of Diffusion ModelsHila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin 等ICLR 2024 · 被引用 38 次
- Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsMor Geva, Jasmijn Bastings, Katja Filippova, Amir GlobersonEMNLP 2023 · 被引用 38 次
- Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalGuangyuan Ma, Xing Wu, Zijia Lin, Songlin HuSIGIR 2024 · 被引用 5 次
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 被引用 5 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni 等NeurIPS 2022 · 被引用 506 次
相关 Paper
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- LED: Lexicon-Enlightened Dense Retriever for Large-Scale RetrievalKai Zhang, Chongyang Tao, Tao Shen, Can Xu 等WWW 2023 · 被引用 27 次
- Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense RetrievalChaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao 等ACL 2024 · 被引用 10 次
- BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense RetrievalShicheng Xu, Liang Pang, Huawei Shen, Xueqi ChengACL 2023 · 被引用 8 次
