F2-Softmax: Diversifying Neural Text Generation via Frequency Factorized Softmax
Byung-Ju Choi, Jimin Hong, David Keetae Park, Sang Wan Lee
摘要
Despite recent advances in neural text generation, encoding the rich diversity in human language remains elusive. We argue that the sub-optimal text generation is mainly attributable to the imbalanced token distribution, which particularly misdirects the learning model when trained with the maximumlikelihood objective. As a simple yet effective remedy, we propose two novel methods, F 2 -Softmax and MefMax, for a balanced training even with the skewed frequency distribution. MefMax assigns tokens uniquely to frequency classes, trying to group tokens with similar frequencies and equalize frequency mass between the classes. F 2 -Softmax then decomposes a probability distribution of the target token into a product of two conditional probabilities of (i) frequency class, and (ii) token from the target frequency class. Models learn more uniform probability distributions because they are confined to subsets of vocabularies. Significant performance gains on seven relevant metrics suggest the supremacy of our approach in improving not only the diversity but also the quality of generated texts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Focus Attention: Promoting Faithfulness and Diversity in SummarizationRahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe 等ACL 2021
- Diversifying Dialog Generation via Adaptive Label SmoothingYida Wang, Yinhe Zheng, Yong Jiang, Minlie HuangACL 2021
它引用的顶会 Paper2
相关 Paper
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng 等EMNLP 2020 · 被引用 32 次
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationXiang Lin, Simeng Han, Shafiq R. JotyICML 2021 · 被引用 30 次
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceHaojin Wang, Zining Zhu, Freda ShiEMNLP 2025
- Data-dependent Gaussian Prior Objective for Language GenerationZuchao Li, Rui Wang, Kehai Chen, Masao Utiyama 等ICLR 2020 · 被引用 68 次
