Learning Probabilistic Box Embeddings for Effective and Efficient Ranking
Lang Mei, Jiaxin Mao, Gang Guo, Ji-Rong Wen
Abstract
Ranking has been one of the most important tasks in information retrieval. With the development of deep representation learning, many researchers propose to encode both the query and items into embedding vectors and rank the items according to the inner product or distance measures in the embedding space. However, the ranking models based on vector embeddings may have shortages in effectiveness and efficiency. For effectiveness, they lack the intrinsic ability to model the diversity and uncertainty of queries and items in ranking. For efficiency, nearest neighbor search in a large collection of item vectors can be costly. In this work, we propose to use the recently proposed probabilistic box embeddings for effective and efficient ranking, in which queries and items are parameterized as high-dimensional axis-aligned hyper-rectangles. For effectiveness, we utilize probabilistic box embeddings to model the diversity and uncertainty with the overlapping relations of the hyper-rectangles, and prove that such overlapping measure is a kernel function which can be adopted in other kernel-based methods. For efficiency, we propose a box embedding-based indexing method, which can safely filter irrelevant items and reduce the retrieval latency. We further design a training strategy to increase the proportion of irrelevant items that can be filtered by the index. Experiments on public datasets show that the box embeddings and the box embedding-based indexing approaches are effective and efficient in two ranking tasks: ad hoc retrieval and product recommendation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff42e86b-80e5-4cec-8731-07ac380daa93Cited by top-tier papers2
- BoxLM: Unifying Structures and Semantics of Medical Concepts for Diagnosis Prediction in HealthcareYanchao Tan, Hang Lv, Yunfei Zhan, Guofang Ma et al.ICML 2025
- A Geometric Approach to Personalized Recommendation with Set-Theoretic Constraints Using Box EmbeddingsShib Sankar Dasgupta, Michael Boratko, Andrew McCallumICML 2025
Builds on5
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Query2box: Reasoning over Knowledge Graphs in Vector Space Using Box EmbeddingsHongyu Ren, Weihua Hu, Jure LeskovecICLR 2020 · 355 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Improving Local Identifiability in Probabilistic Box EmbeddingsShib Sankar Dasgupta, Michael Boratko, Dongxu Zhang, Luke Vilnis et al.NeurIPS 2020 · 75 citations
Related papers
- A Single Vector Is Not Enough: Taxonomy Expansion via Box EmbeddingsSong Jiang, Qiyue Yao, Qifan Wang, Yizhou SunWWW 2023 · 20 citations
- Enhancing Recommendation Accuracy and Diversity with Box Embedding: A Universal FrameworkCheng Wu, Shaoyun Shi, Chaokun Wang, Ziyang Liu et al.WWW 2024 · 9 citations
- BoxCD: Leveraging Contrastive Probabilistic Box Embedding for Effective and Efficient Learner ModelingWeibo Gao, Qi Liu, Linan Yue, Fangzhou Yao et al.WWW 2025 · 3 citations
- When Box Meets Graph Neural Network in Tag-aware RecommendationFake Lin, Ziwei Zhao, Xi Zhu, Da Zhang et al.KDD 2024 · 6 citations
- Optimizing Probabilistic Box Embeddings with Distance MeasuresLang Mei, Jiaxin Mao, Ji-Rong WenICDE 2024 · 1 citation
