Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe
Chong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka, Felix X. Yu, Sanjiv Kumar
Abstract
Dual encoder (DE) models, where a pair of matching query and document are embedded into similar vector representations, are widely used in information retrieval due to their simplicity and scalability. However, the Euclidean geometry of the embedding space limits the expressive power of DEs, which may compromise their quality. This paper investigates such limitations in the context of hierarchical retrieval (HR), where the document set has a hierarchical structure and the matching documents for a query are all of its ancestors. We first prove that DEs are feasible for HR as long as the embedding dimension is linear in the depth of the hierarchy and logarithmic in the number of documents. Then we study the problem of learning such embeddings in a standard retrieval setup where DEs are trained on samples of matching query and document pairs. Our experiments reveal a lost-in-the-long-distance phenomenon, where retrieval accuracy degrades for documents further away in the hierarchy. To address this, we introduce a pretrain-finetune recipe that significantly improves long-distance retrieval without sacrificing performance on closer documents. We experiment on a realistic hierarchy from WordNet for retrieving documents at various levels of abstraction, and show that pretrain-finetune boosts the recall on long-distance pairs from 19% to 76%. Finally, we demonstrate that our method improves retrieval of relevant products on a shopping queries dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dc62f66-2a03-41bf-8d27-5d25cd90c83eCited by top-tier papers1
Ask how each one uses itBuilds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
- Language Models as Hierarchy EncodersYuan He, Moy Yuan, Jiaoyan Chen, Ian HorrocksNeurIPS 2024 · 37 citations
- TPU-KNN: K Nearest Neighbor Search at Peak FLOP/sFelix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo et al.NeurIPS 2022 · 31 citations
Related papers
- In defense of dual-encoders for neural rankingAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim et al.ICML 2022 · 29 citations
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM EmbeddingsXueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins et al.ACL 2026 · 3 citations
- Graph-based Hierarchical Relevance Matching Signals for Ad-hoc RetrievalXueli Yu, Weizhi Xu, Zeyu Cui, Shu Wu et al.WWW 2021 · 20 citations
- Structure and Semantics Preserving Document RepresentationsNatraj Raman, Sameena Shah, Manuela VelosoSIGIR 2022 · 5 citations
- Wikiformer: Pre-training with Structured Information of Wikipedia for Ad-Hoc RetrievalWeihang Su, Qingyao Ai, Xiangsheng Li, Jia Chen et al.AAAI 2024
