ReadOnce Transformers: Reusable Representations of Text for Transformers
Shih-Ting Lin, Ashish Sabharwal, Tushar Khot
摘要
We present READONCE Transformers, an approach to convert a transformer-based model into one that can build an informationcapturing, task-independent, and compressed representation of text. The resulting representation is reusable across different examples and tasks, thereby requiring a document shared across many examples or tasks to only be read once. This leads to faster training and evaluation of models. Additionally, we extend standard text-to-text transformer models to Representation+Text-to-text models, and evaluate on multiple downstream tasks: multihop QA, abstractive QA, and long-document summarization. Our one-time computed representation results in a 2x-5x speedup compared to standard text-to-text models, while the compression also allows existing language models to handle longer documents without the need for designing new pre-trained models. 1 * The author's work was primarily done during an internship at the Allen Institute for AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- PlotMachines: Outline-Conditioned Generation with Dynamic Plot State TrackingHannah Rashkin, Asli Celikyilmaz, Yejin Choi, Jianfeng GaoEMNLP 2020 · 被引用 100 次
相关 Paper
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 被引用 273 次
- DeFormer: Decomposing Pre-trained Transformers for Faster Question AnsweringQingqing Cao, Harsh Trivedi, Aruna Balasubramanian, Niranjan BalasubramanianACL 2020 · 被引用 61 次
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser 等NeurIPS 2021 · 被引用 127 次
- Dodo: Dynamic Contextual Compression for Decoder-only LMsGuanghui Qin, Corby Rosset, Ethan C. Chau, Nikhil Rao 等ACL 2024
- Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsSean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto 等SIGIR 2020 · 被引用 62 次
