Copy is All You Need
Tian Lan, Deng Cai, Yan Wang, Heyan Huang, Xian-Ling Mao
摘要
The dominant text generation models compose the output by sequentially selecting words from a fixed vocabulary. In this paper, we formulate text generation as progressively copying text segments (e.g., words or phrases) from an existing text collection. We compute the contextualized representations of meaningful text segments and index them using efficient vector search toolkits. The task of text generation is then decomposed into a series of copy-and-paste operations: at each time step, we seek suitable text spans from the text collection rather than selecting from a standalone vocabulary. Experiments on the standard language modeling benchmark (WikiText-103) show that our approach achieves better generation quality according to both automatic and human evaluations. Besides, its inference efficiency is comparable to token-level autoregressive models thanks to the reduction of decoding steps. We also show that our approach allows for effective domain adaptation by simply switching to domain-specific text collection without extra training. Finally, we observe that our approach attains additional performance gains by simply scaling up to larger text collections, again without further training. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- SILO Language Models: Isolating Legal Risk In a Nonparametric DatastoreSewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi 等ICLR 2024 · 被引用 91 次
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen 等ACL 2026 · 被引用 57 次
- Thought Propagation: an Analogical Approach to Complex Reasoning with Large Language ModelsJunchi Yu, Ran He, Zhitao YingICLR 2024 · 被引用 44 次
- Nearest Neighbor Speculative Decoding for LLM Generation and AttributionMinghan Li, Xilun Chen, Ari Holtzman, Beidi Chen 等NeurIPS 2024 · 被引用 29 次
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionJinghui Lu, Yanjie Wang, Ziwei Yang, Xuejing Liu 等NeurIPS 2024 · 被引用 22 次
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
相关 Paper
- GNN-LM: Language Modeling based on Global Contexts via GNNYuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun 等ICLR 2022 · 被引用 46 次
- Mixture of Inputs: Text Generation Beyond Discrete Token SamplingYufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang 等NeurIPS 2025
- Iterative GNN-based Decoder for Question GenerationZichu Fei, Qi Zhang, Yaqian ZhouEMNLP 2021 · 被引用 1 次
- Copy That! Editing Sequences by Copying SpansSheena Panthaplackel, Miltiadis Allamanis, Marc BrockschmidtAAAI 2021 · 被引用 28 次
- Nugget: Neural Agglomerative Embeddings of TextGuanghui Qin, Benjamin Van DurmeICML 2023 · 被引用 24 次
