Progressively Optimized Bi-Granular Document Representation for Scalable Embedding Based Retrieval
Shitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang, Yingxia Shao, Defu Lian, Chaozhuo Li, Hao Sun, Denvy Deng, Liangjie Zhang, Qi Zhang, Xing Xie
摘要
Ad-hoc search calls for the selection of appropriate answers from a massive-scale corpus. Nowadays, the embedding-based retrieval (EBR) becomes a promising solution, where deep learning based document representation and ANN search techniques are allied to handle this task. However, a major challenge is that the ANN index can be too large to fit into memory, given the considerable size of answer corpus. In this work, we tackle this problem with Bi-Granular Document Representation, where the lightweight sparse embeddings are indexed and standby in memory for coarse-grained candidate search, and the heavyweight dense embeddings are hosted in disk for fine-grained post verification. For the best of retrieval accuracy, a Progressive Optimization framework is designed. The sparse embeddings are learned ahead for high-quality search of candidates. Conditioned on the candidate distribution induced by the sparse embeddings, the dense embeddings are continuously learned to optimize the discrimination of ground-truth from the shortlisted candidates. Besides, two techniques: the contrastive quantization and the locality-centric sampling are introduced for the learning of sparse and dense embeddings, which substantially contribute to their performances. Thanks to the above features, our method effectively handles massive-scale EBR with strong advantages in accuracy: with up to +4.3% recall gain on million-scale corpus, and up to +17.5% recall gain on billion-scale corpus. Besides, Our method is applied to a major sponsored search platform with substantial gains on revenue (+1.95%), Recall (+1.01%) and CTR (+0.49%). Our code is available at https://github.com/microsoft/BiDR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- VBASE: Unifying Online Vector Similarity Search and Relational Queries via Relaxed MonotonicityQianxi Zhang, Shuotao Xu, Qi Chen, Guoxin Sui 等OSDI 2023 · 被引用 75 次
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang 等SIGIR 2022 · 被引用 31 次
- Tribase: A Vector Data Query Engine for Reliable and Lossless Pruning Compression using Triangle InequalitiesQian Xu, Juan Yang, Feng Zhang, Junda Pan 等SIGMOD 2025 · 被引用 14 次
- LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN SearchXimu Zeng, Liwei Deng, Penghao Chen, Xu Chen 等WWW 2025 · 被引用 10 次
- Efficient Dynamic Indexing for Range Filtered Approximate Nearest Neighbor SearchFangyuan Zhang, Mengxu Jiang, Guanhao Hou, Jieming Shi 等SIGMOD 2025 · 被引用 8 次
它引用的顶会 Paper11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang 等ICLR 2020 · 被引用 325 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
相关 Paper
- Multi-View Document Representation Learning for Open-Domain Dense RetrievalShunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang 等ACL 2022 · 被引用 80 次
- Efficient Inverted Indexes for Approximate Retrieval over Learned Sparse RepresentationsSebastian Bruch, Franco Maria Nardini, Cosimo Rulli, Rossano VenturiniSIGIR 2024 · 被引用 47 次
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang 等ACL 2021
- Improving the Accuracy of Dense Retrieval on the Quantized Indexes via Gradient Optimization of the Target EmbeddingsCong Tan, Yongqi Shao, Hong Huo, Tao FangAAAI 2026
- COCO-DR: Combating the Distribution Shift in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust LearningYue Yu, Chenyan Xiong, Si Sun, Chao Zhang 等EMNLP 2022 · 被引用 21 次
