Optimizing Retrieval-augmented Reader Models via Token Elimination
Moshe Berchansky, Peter Izsak, Avi Caciularu, Ido Dagan, Moshe Wasserblat
摘要
Fusion-in-Decoder (FiD) is an effective retrieval-augmented language model applied across a variety of open-domain tasks, such as question answering, fact checking, etc. In FiD, supporting passages are first retrieved and then processed using a generative model (Reader), which can cause a significant bottleneck in decoding time, particularly with long outputs. In this work, we analyze the contribution and necessity of all the retrieved passages to the performance of reader models, and propose eliminating some of the retrieved information, at the token level, that might not contribute essential information to the answer generation process. We demonstrate that our method can reduce run-time by up to 62.2%, with only a 2% reduction in performance, and in some cases, even improve the performance results. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Leveraging Attention to Effectively Compress Prompts for Long-Context LLMsYunlong Zhao, Haoran Wu, Bo XuAAAI 2025 · 被引用 10 次
- Transformers are Multi-State RNNsMatanel Oren, Michael Hassid, Yarden Nir, Yossi Adi 等EMNLP 2024 · 被引用 8 次
- Rethinking Token Reduction for State Space ModelsZheng Zhan, Yushu Wu, Zhenglun Kong, Changdi Yang 等EMNLP 2024 · 被引用 4 次
- RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation SystemsShaobo Li, Yirui Zhou, Yuan Xu, Kevin Chen 等VLDB 2026 · 被引用 3 次
- Preference-Guided Refactored Tuning for Retrieval Augmented Code GenerationXinyu Gao, Yun Xiong, Deze Wang, Zhenhan Guan 等ASE 2024 · 被引用 1 次
它引用的顶会 Paper14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 被引用 264 次
相关 Paper
- FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence SelectionYufei Huang, Xu Han, Maosong SunACL 2024
- KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question AnsweringDonghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu 等ACL 2022 · 被引用 108 次
- Modeling Multi-hop Question Answering as Single Sequence PredictionSemih Yavuz, Kazuma Hashimoto, Yingbo Zhou, Nitish Shirish Keskar 等ACL 2022 · 被引用 35 次
- REANO: Optimising Retrieval-Augmented Reader Models through Knowledge Graph GenerationJinyuan Fang, Zaiqiao Meng, Craig MacDonaldACL 2024 · 被引用 7 次
- Pre-computed memory or on-the-fly encoding? A hybrid approach to retrieval augmentation makes the most of your computeMichiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Joshua Ainslie 等ICML 2023 · 被引用 20 次
