Learning Sparse Prototypes for Text Generation
Junxian He, Taylor Berg-Kirkpatrick, Graham Neubig
Abstract
Prototype-driven text generation uses non-parametric models that first choose from a library of sentence "prototypes" and then modify the prototype to generate the output text. While effective, these methods are inefficient at test time as a result of needing to store and index the entire training corpus. Further, existing methods often require heuristics to identify which prototypes to reference at training time. In this paper, we propose a novel generative model that automatically learns a sparse prototype support set that, nonetheless, achieves strong language modeling performance. This is achieved by (1) imposing a sparsity-inducing prior on the prototype selection distribution, and (2) utilizing amortized variational inference to learn a prototype retrieval function. In experiments, our model outperforms previous prototype-driven language models while achieving up to a 1000x memory reduction, as well as a 1000x speed-up at test time. More interestingly, we show that the learned prototypes are able to capture semantics and syntax at different granularity as we vary the sparsity of prototype selection, and that certain sentence attributes can be controlled by specifying the prototype for generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c1b2592-e77c-4a55-adea-a4be2cfa95adCited by top-tier papers9
- Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalUri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta et al.ICML 2022 · 79 citations
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani et al.EMNLP 2022 · 59 citations
- Why do Nearest Neighbor Language Models Work?Frank F. Xu, Uri Alon, Graham NeubigICML 2023 · 33 citations
- Controllable Semantic Parsing via Retrieval AugmentationPanupong Pasupat, Yuan Zhang, Kelvin GuuEMNLP 2021 · 29 citations
- Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsKalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta et al.ACL 2022 · 28 citations
Builds on2
Related papers
- Efficient Nearest Neighbor Language ModelsJunxian He, Graham Neubig, Taylor Berg-KirkpatrickEMNLP 2021
- AUTOSUMM: Automatic Model Creation for Text SummarizationSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee et al.EMNLP 2021 · 1 citation
- Discourse-Aware Soft Prompting for Text GenerationMarjan Ghazvininejad, Vladimir Karpukhin, Vera Gor, Asli CelikyilmazEMNLP 2022 · 6 citations
- Building, Reusing, and Generalizing Abstract Representations from Concrete SequencesShuchen Wu, Mirko Thalmann, Peter Dayan, Zeynep Akata et al.ICLR 2025
- Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsXueliang Zhao, Wei Wu, Can Xu, Chongyang Tao et al.EMNLP 2020 · 153 citations
