Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, Allan Hanbury
摘要
A vital step towards the widespread adoption of neural retrieval models is their resource efficiency throughout the training, indexing and query workflows. The neural IR community made great advancements in training effective dual-encoder dense retrieval (DR) models recently. A dense text retrieval model uses a single vector representation per query and passage to score a match, which enables low-latency first-stage retrieval with a nearest neighbor search. Increasingly common, training approaches require enormous compute power, as they either conduct negative passage sampling out of a continuously updating refreshing index or require very large batch sizes. Instead of relying on more compute capability, we introduce an efficient topic-aware query and balanced margin sampling technique, called TAS-Balanced. We cluster queries once before training and sample queries out of a cluster per batch. We train our lightweight 6-layer DR model with a novel dual-teacher supervision that combines pairwise and in-batch negative teachers. Our method is trainable on a single consumer-grade GPU in under 48 hours. We show that our TAS-Balanced training method achieves state-of-the-art low-latency (64ms per query) results on two TREC Deep Learning Track query sets. Evaluated on [email protected], we outperform BM25 by 44%, a plainly trained DR by 19%, docT5query by 11%, and the previous best DR model by 5%. Additionally, TAS-Balanced produces the first dense retriever that outperforms every other method on recall at any cutoff on TREC-DL and allows more resource intensive re-ranking models to operate on fewer passages to improve results further.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper69
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 被引用 211 次
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang 等NeurIPS 2023 · 被引用 151 次
- RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-rankingRuiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao 等EMNLP 2021 · 被引用 147 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- Scalable and Effective Generative Information RetrievalHansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar 等WWW 2024 · 被引用 72 次
它引用的顶会 Paper8
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel 等NeurIPS 2020 · 被引用 805 次
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
相关 Paper
- Improving the Accuracy of Dense Retrieval on the Quantized Indexes via Gradient Optimization of the Target EmbeddingsCong Tan, Yongqi Shao, Hong Huo, Tao FangAAAI 2026
- Constructing Tree-based Index for Efficient and Effective Dense RetrievalHaitao Li, Qingyao Ai, Jingtao Zhan, Jiaxin Mao 等SIGIR 2023 · 被引用 21 次
- TITE: Token-Independent Text Encoder for Information RetrievalFerdinand Schlatt, Tim Hagen, Martin Potthast, Matthias HagenSIGIR 2025 · 被引用 2 次
- Adversarial Retriever-Ranker for Dense Text RetrievalHang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv 等ICLR 2022 · 被引用 137 次
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin 等AAAI 2023 · 被引用 34 次
