Multivariate Representation Learning for Information Retrieval
Hamed Zamani, Michael Bendersky
Abstract
Dense retrieval models use bi-encoder network architectures for learning query and document representations. These representations are often in the form of a vector representation and their similarities are often computed using the dot product function. In this paper, we propose a new representation learning framework for dense retrieval. Instead of learning a vector for each query and document, our framework learns a multivariate distribution and uses negative multivariate KL divergence to compute the similarity between distributions. For simplicity and efficiency reasons, we assume that the distributions are multivariate normals and then train large language models to produce mean and variance vectors for these distributions. We provide a theoretical foundation for the proposed framework and show that it can be seamlessly integrated into the existing approximate nearest neighbor algorithms to perform retrieval efficiently. We conduct an extensive suite of experiments on a wide range of datasets, and demonstrate significant improvements compared to competitive dense retrieval models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40a1ae50-c20f-4754-8143-cec88e4aed78Cited by top-tier papers2
- Bridging Personalization and Control in Scientific Personalized SearchSheshera Mysore, Garima Dhanania, Kishor Patil, Surya Kallumadi et al.SIGIR 2025 · 1 citation
- Efficient Partition-based Approaches for Diversified Top-k Subgraph MatchingLiuyi Chen, Yuchen Hu, Zhengyi Yang, Xu Zhou et al.VLDB 2026
Builds on7
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipJunfeng Kang, Rui Li, Qi Liu, Zhenya Huang et al.AAAI 2025 · 2 citations
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang et al.ACL 2022 · 80 citations
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang et al.ACL 2021
- Generative Retrieval as Multi-Vector Dense RetrievalShiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen et al.SIGIR 2024 · 14 citations
- Learning Retrieval Models with Sparse AutoencodersThibault Formal, Maxime Louis, Hervé Déjean, Stéphane ClinchantICLR 2026 · 9 citations
