Distilling Large Embeddings via Hyperspherical Householder Quantization
Yihang Wang, Bin Wu, Yueyang Su, Tianfu Zhang, Yiqi Du, Lei Yu, Jiafeng Guo, Xueqi Cheng
Abstract
Large embedding models have become the backbone of modern retrieval systems, offering strong semantic representations at the cost of substantial storage and computation. While recent work explores quantizing embeddings into discrete document identifiers for generative retrieval, most existing approaches rely on Euclidean quantization, which is poorly aligned with the angular geometry induced by contrastive embedding training and often requires long identifier sequences to preserve semantic fidelity. In this work, we propose Hyperspherical Householder Quantization (HHQ), a geometry-aware distillation method that compresses large embeddings into short discrete representations via iterative Householder transformations on the unit hypersphere. By explicitly preserving cosine similarity at each step, HHQ distills semantic structure into compact identifiers that remain faithful to the original embedding space. To support reliable generation of these identifiers, we introduce constrained supervised fine-tuning and tree-aware dynamic masking to enforce structural validity during training and inference. Experiments on NQ and MS MARCO show that HHQ achieves competitive or superior retrieval performance using only five tokens per document, substantially reducing decoding cost while retaining strong semantic retrieval accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20081495-0e7f-4760-a1ab-4132de72e4daBuilds on11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao et al.NeurIPS 2022 · 242 citations
Related papers
- Weakly Supervised Deep Hyperspherical Quantization for Image RetrievalJinpeng Wang, Bin Chen, Qiang Zhang, Zaiqiao Meng et al.AAAI 2021 · 13 citations
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang et al.SIGIR 2022 · 31 citations
- Hyperbolic RQ-VAE enhanced Generative Recommendation with Differential-Length Codebook StrategyAoran Zhang, Yu-Bin Yang, Yonghong YuICML 2026
- HiHPQ: Hierarchical Hyperbolic Product Quantization for Unsupervised Image RetrievalZexuan Qiu, Jiahong Liu, Yankai Chen, Irwin KingAAAI 2024 · 14 citations
- Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product QuantizationZexuan Qiu, Qinliang Su, Jianxing Yu, Shijing SiEMNLP 2022 · 4 citations
