Improving Text Generation with Student-Forcing Optimal Transport
Jianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuh-Chen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin
摘要
Neural language models are often trained with maximum likelihood estimation (MLE), where the next word is generated conditioned on the ground-truth word tokens. During testing, however, the model is instead conditioned on previously generated tokens, resulting in what is termed exposure bias. To reduce this gap between training and testing, we propose using optimal transport (OT) to match the sequences generated in these two modes. An extension is further proposed to improve the OT learning, based on the structural and contextual information of the text sequences. The effectiveness of the proposed method is validated on machine translation, text summarization, and text generation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
- Toward Interpretable Semantic Textual Similarity via Optimal Transport-based Contrastive Sentence LearningSeonghyeon Lee, Dongha Lee, Seongbo Jang, Hwanjo YuACL 2022 · 被引用 25 次
- Re-evaluating Word Mover's DistanceRyoma Sato, Makoto Yamada, Hisashi KashimaICML 2022 · 被引用 25 次
- Event Causality Extraction via Implicit Cause-Effect InteractionsJintao Liu, Zequn Zhang, Kaiwen Wei, Zhi Guo 等EMNLP 2023 · 被引用 8 次
- OTSeq2Set: An Optimal Transport Enhanced Sequence-to-Set Model for Extreme Multi-label Text ClassificationJie Cao, Yin ZhangEMNLP 2022 · 被引用 6 次
它引用的顶会 Paper2
相关 Paper
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
- Scheduled Sampling Based on Decoding Steps for Neural Machine TranslationYijin Liu, Fandong Meng, Yufeng Chen, Jinan Xu 等EMNLP 2021 · 被引用 9 次
- Sequence Generation with Optimal-Transport-Enhanced Reinforcement LearningLiqun Chen, Ke Bai, Chenyang Tao, Yizhe Zhang 等AAAI 2020 · 被引用 14 次
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 被引用 9 次
- ColdGANs: Taming Language GANs with Cautious Sampling StrategiesThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 等NeurIPS 2020 · 被引用 19 次
