Sparse Text Generation
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
Abstract
Current state-of-the-art text generators build on powerful language models such as GPT-2, achieving impressive performance. However, to avoid degenerate text, they require sampling from a modified softmax, via temperature parameters or ad-hoc truncation techniques, as in top-k or nucleus sampling. This creates a mismatch between training and testing conditions. In this paper, we use the recently introduced entmax transformation to train and sample from a natively sparse language model, avoiding this mismatch. The result is a text generator with favorable performance in terms of fluency and consistency, fewer repetitions, and n-gram diversity closer to human text. In order to evaluate our model, we propose three new metrics for comparing sparse or truncated distributions: -perplexity, sparsemax score, and Jensen-Shannon divergence. Human-evaluated experiments in story completion and dialogue generation show that entmax sampling leads to more engaging and coherent stories and conversations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 879316f2-d83d-4c90-8977-06ad2505f64aCited by top-tier papers11
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi et al.ACL 2023 · 35 citations
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong et al.ACL 2023 · 12 citations
- Frictional Agent Alignment Framework: Slow Down and Don't Break ThingsAbhijnan Nath, Carine Graff, Andrei Bachinin, Nikhil KrishnaswamyACL 2025 · 8 citations
- Language Model Decoding as Direct Metrics OptimizationHaozhe Ji, Pei Ke, Hongning Wang, Minlie HuangICLR 2024 · 8 citations
Builds on3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck et al.ACL 2020 · 120 citations
Related papers
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 19 citations
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
- F2-Softmax: Diversifying Neural Text Generation via Frequency Factorized SoftmaxByung-Ju Choi, Jimin Hong, David Keetae Park, Sang Wan LeeEMNLP 2020 · 16 citations
- TS: Training with Sparsemax+, Testing with Softmax for Accurate and Diverse LLM Fine-TuningZiyang Xu, Ananthu Rajendran Pillai, Yinghua Yao, Yuangang PanICLR 2026
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 3 citations
