Reduce Catastrophic Forgetting of Dense Retrieval Training with Teleportation Negatives
Si Sun, Chenyan Xiong, Yue Yu, Arnold Overwijk, Zhiyuan Liu, Jie Bao
摘要
In this paper, we investigate the instability in the standard dense retrieval training, which iterates between model training and hard negative selection using the being-trained model. We show the catastrophic forgetting phenomena behind the training instability, where models learn and forget different negative groups during training iterations. We then propose ANCE-Tele, which accumulates momentum negatives from past iterations and approximates future iterations using lookahead negatives, as “teleportations” along the time axis to smooth the learning process. On web search and OpenQA, ANCE-Tele outperforms previous state-of-the-art systems of similar size, eliminates the dependency on sparse retrieval negatives, and is competitive among systems using significantly more (50x) parameters. Our analysis demonstrates that teleportation negatives reduce catastrophic forgetting and improve convergence speed for dense retrieval training. The source code of this paper is available at https://github.com/OpenMatch/ANCE-Tele.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TriSampler: A Better Negative Sampling Principle for Dense RetrievalZhen Yang, Zhou Shao, Yuxiao Dong, Jie TangAAAI 2024 · 被引用 17 次
- HOBIT: Hardness Optimized Batch Sampling for InfoNCE TrainingHimanshu Dutta, Lokesh Nagalapatti, Yashoteja PrabhuICML 2026
它引用的顶会 Paper15
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo 等SIGIR 2021 · 被引用 242 次
相关 Paper
- A Fresh Take on Stale Embeddings: Improving Dense Retriever Training with Corrector NetworksNicholas Monath, Will Sussman Grathwohl, Michael Boratko, Rob Fergus 等ICML 2024 · 被引用 1 次
- PROD: Progressive Distillation for Dense RetrievalZhenghao Lin, Yeyun Gong, Xiao Liu, Hang Zhang 等WWW 2023 · 被引用 33 次
- Hot-Refresh Model Upgrades with Regression-Free Compatible Training in Image RetrievalBinjie Zhang, Yixiao Ge, Yantao Shen, Yu Li 等ICLR 2022 · 被引用 13 次
- A Gradient Accumulation Method for Dense Retriever under Memory ConstraintJaehee Kim, Yukyung Lee, Pilsung KangNeurIPS 2024 · 被引用 10 次
- HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval GeneralizationZefeng Cai, Chongyang Tao, Tao Shen, Can Xu 等ICLR 2023
