Learning Sparse Prototypes for Text Generation
Junxian He, Taylor Berg-Kirkpatrick, Graham Neubig
摘要
Prototype-driven text generation uses non-parametric models that first choose from a library of sentence "prototypes" and then modify the prototype to generate the output text. While effective, these methods are inefficient at test time as a result of needing to store and index the entire training corpus. Further, existing methods often require heuristics to identify which prototypes to reference at training time. In this paper, we propose a novel generative model that automatically learns a sparse prototype support set that, nonetheless, achieves strong language modeling performance. This is achieved by (1) imposing a sparsity-inducing prior on the prototype selection distribution, and (2) utilizing amortized variational inference to learn a prototype retrieval function. In experiments, our model outperforms previous prototype-driven language models while achieving up to a 1000x memory reduction, as well as a 1000x speed-up at test time. More interestingly, we show that the learned prototypes are able to capture semantics and syntax at different granularity as we vary the sparsity of prototype selection, and that certain sentence attributes can be controlled by specifying the prototype for generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalUri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta 等ICML 2022 · 被引用 79 次
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani 等EMNLP 2022 · 被引用 59 次
- Why do Nearest Neighbor Language Models Work?Frank F. Xu, Uri Alon, Graham NeubigICML 2023 · 被引用 33 次
- Controllable Semantic Parsing via Retrieval AugmentationPanupong Pasupat, Yuan Zhang, Kelvin GuuEMNLP 2021 · 被引用 29 次
- Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsKalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta 等ACL 2022 · 被引用 28 次
它引用的顶会 Paper2
相关 Paper
- Efficient Nearest Neighbor Language ModelsJunxian He, Graham Neubig, Taylor Berg-KirkpatrickEMNLP 2021
- AUTOSUMM: Automatic Model Creation for Text SummarizationSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee 等EMNLP 2021 · 被引用 1 次
- Discourse-Aware Soft Prompting for Text GenerationMarjan Ghazvininejad, Vladimir Karpukhin, Vera Gor, Asli CelikyilmazEMNLP 2022 · 被引用 6 次
- Building, Reusing, and Generalizing Abstract Representations from Concrete SequencesShuchen Wu, Mirko Thalmann, Peter Dayan, Zeynep Akata 等ICLR 2025
- Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsXueliang Zhao, Wei Wu, Can Xu, Chongyang Tao 等EMNLP 2020 · 被引用 153 次
