Sparse Text Generation
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
摘要
Current state-of-the-art text generators build on powerful language models such as GPT-2, achieving impressive performance. However, to avoid degenerate text, they require sampling from a modified softmax, via temperature parameters or ad-hoc truncation techniques, as in top-k or nucleus sampling. This creates a mismatch between training and testing conditions. In this paper, we use the recently introduced entmax transformation to train and sample from a natively sparse language model, avoiding this mismatch. The result is a text generator with favorable performance in terms of fluency and consistency, fewer repetitions, and n-gram diversity closer to human text. In order to evaluate our model, we propose three new metrics for comparing sparse or truncated distributions: -perplexity, sparsemax score, and Jensen-Shannon divergence. Human-evaluated experiments in story completion and dialogue generation show that entmax sampling leads to more engaging and coherent stories and conversations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi 等ACL 2023 · 被引用 35 次
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong 等ACL 2023 · 被引用 12 次
- Frictional Agent Alignment Framework: Slow Down and Don't Break ThingsAbhijnan Nath, Carine Graff, Andrei Bachinin, Nikhil KrishnaswamyACL 2025 · 被引用 8 次
- Language Model Decoding as Direct Metrics OptimizationHaozhe Ji, Pei Ke, Hongning Wang, Minlie HuangICLR 2024 · 被引用 8 次
它引用的顶会 Paper3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck 等ACL 2020 · 被引用 120 次
相关 Paper
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 被引用 19 次
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta 等ICLR 2024 · 被引用 31 次
- F2-Softmax: Diversifying Neural Text Generation via Frequency Factorized SoftmaxByung-Ju Choi, Jimin Hong, David Keetae Park, Sang Wan LeeEMNLP 2020 · 被引用 16 次
- TS: Training with Sparsemax+, Testing with Softmax for Accurate and Diverse LLM Fine-TuningZiyang Xu, Ananthu Rajendran Pillai, Yinghua Yao, Yuangang PanICLR 2026
- Sparse Communication via Mixed DistributionsAntónio Farinhas, Wilker Aziz, Vlad Niculae, André F. T. MartinsICLR 2022 · 被引用 3 次
