Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang, Hai Zhao
摘要
For typical sequence prediction problems like language generation, maximum likelihood estimation (MLE) has been commonly adopted as it encourages the predicted sequence most consistent with the ground-truth sequence to have the highest probability of occurring. However, MLE focuses on a once-for-all matching between the predicted sequence and gold-standard consequently, treating all incorrect predictions as being equally incorrect. We call such a drawback negative diversity ignorance in this paper. Treating all incorrect predictions as equal unfairly downplays the nuance of these sequences' detailed token-wise structure. To counteract this, we augment the MLE loss by introducing an extra KL divergence term which is derived from comparing a data-dependent Gaussian prior and the detailed training prediction. The proposed data-dependent Gaussian prior objective (D2GPo) is defined over a prior topological order of tokens, poles apart from the data-independent Gaussian prior (L2 regularization) commonly adopted for smoothing the training of MLE. Experimental results show that the proposed method can effectively make use of more detailed prior in the data and significantly improve the performance of typical language generation tasks, including supervised and unsupervised machine translation, text summarization, storytelling, and image caption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive PriorSang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 等ICLR 2022 · 被引用 117 次
- Cross-lingual Retrieval for Iterative Self-Supervised TrainingChau Tran, Yuqing Tang, Xian Li, Jiatao GuNeurIPS 2020 · 被引用 76 次
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou 等AAAI 2021 · 被引用 57 次
- Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine TranslationSongming Zhang, Yijin Liu, Fandong Meng, Yufeng Chen 等ACL 2022 · 被引用 13 次
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu 等EMNLP 2020 · 被引用 11 次
- Tailoring Language Generation Models under Total Variation DistanceHaozhe Ji, Pei Ke, Zhipeng Hu, Rongsheng Zhang 等ICLR 2023 · 被引用 2 次
- F2-Softmax: Diversifying Neural Text Generation via Frequency Factorized SoftmaxByung-Ju Choi, Jimin Hong, David Keetae Park, Sang Wan LeeEMNLP 2020 · 被引用 16 次
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen 等ICML 2026
- Adaptive Prior-Dependent Correction Enhanced Reinforcement Learning for Natural Language GenerationWei Cheng, Ziyan Luo, Qiyue YinAAAI 2021 · 被引用 1 次
