Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, Michael Zeng
Abstract
Retrieval-based methods have been shown to be effective in NLP tasks via introducing external knowledge. However, the indexing and retrieving of large-scale corpora bring considerable computational cost. Surprisingly, we found that REtrieving from the traINing datA (REINA) only can lead to significant gains on multiple NLG and NLU tasks. We retrieve the labeled training instances most similar to the input text and then concatenate them with the input to feed into the model to generate the output. Experimental results show that this simple method can achieve significantly better performance on a variety of NLU and NLG tasks, including summarization, machine translation, language modeling, and question answering tasks. For instance, our proposed method achieved state-of-the-art results on XSum, BigPatent, and CommonsenseQA. Our code is released, https://github.com/microsoft/REINA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f71bbf34-705a-4bc2-b87e-016fa67b6298Cited by top-tier papers34
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenXin Cheng, Xun Wang, Xingxing Zhang, Tao Ge et al.NeurIPS 2024 · 156 citations
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii et al.ICLR 2024 · 78 citations
- Test-Time Training on Nearest Neighbors for Large Language ModelsMoritz Hardt, Yu SunICLR 2024 · 66 citations
- Unsupervised Cross-Task Generalization via Retrieval AugmentationBill Yuchen Lin, Kangmin Tan, Chris Miller, Beiwen Tian et al.NeurIPS 2022 · 53 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
Related papers
- Reusing Pre-Training Data at Test Time is a Compute MultiplierAlex Fang, Thomas Voice, Ruoming Pang, Ludwig Schmidt et al.ICLR 2026 · 4 citations
- Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense ReasoningJunho Kim, Soyeon Bak, Mingyu Lee, Minju Hong et al.EMNLP 2025
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck et al.ACL 2022
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang et al.VLDB 2025 · 48 citations
