Distilling Large Embeddings via Hyperspherical Householder Quantization
Yihang Wang, Bin Wu, Yueyang Su, Tianfu Zhang, Yiqi Du, Lei Yu, Jiafeng Guo, Xueqi Cheng
摘要
Large embedding models have become the backbone of modern retrieval systems, offering strong semantic representations at the cost of substantial storage and computation. While recent work explores quantizing embeddings into discrete document identifiers for generative retrieval, most existing approaches rely on Euclidean quantization, which is poorly aligned with the angular geometry induced by contrastive embedding training and often requires long identifier sequences to preserve semantic fidelity. In this work, we propose Hyperspherical Householder Quantization (HHQ), a geometry-aware distillation method that compresses large embeddings into short discrete representations via iterative Householder transformations on the unit hypersphere. By explicitly preserving cosine similarity at each step, HHQ distills semantic structure into compact identifiers that remain faithful to the original embedding space. To support reliable generation of these identifiers, we introduce constrained supervised fine-tuning and tree-aware dynamic masking to enforce structural validity during training and inference. Experiments on NQ and MS MARCO show that HHQ achieves competitive or superior retrieval performance using only five tokens per document, substantially reducing decoding cost while retaining strong semantic retrieval accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni 等NeurIPS 2022 · 被引用 506 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih 等NeurIPS 2022 · 被引用 242 次
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao 等NeurIPS 2022 · 被引用 242 次
相关 Paper
- Weakly Supervised Deep Hyperspherical Quantization for Image RetrievalJinpeng Wang, Bin Chen, Qiang Zhang, Zaiqiao Meng 等AAAI 2021 · 被引用 13 次
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang 等SIGIR 2022 · 被引用 31 次
- Hyperbolic RQ-VAE enhanced Generative Recommendation with Differential-Length Codebook StrategyAoran Zhang, Yu-Bin Yang, Yonghong YuICML 2026
- HiHPQ: Hierarchical Hyperbolic Product Quantization for Unsupervised Image RetrievalZexuan Qiu, Jiahong Liu, Yankai Chen, Irwin KingAAAI 2024 · 被引用 14 次
- Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product QuantizationZexuan Qiu, Qinliang Su, Jianxing Yu, Shijing SiEMNLP 2022 · 被引用 4 次
