MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
Shiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu, Mohit Bansal, Mark Dredze, David S. Rosenberg
摘要
Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q θ relative to the data distribution Pthat is, minimizing the forward cross-entropy, which is equivalent to maximum likelihood estimation (MLE). We have observed that models trained in this way may "over-generalize", in the sense that they produce non-human-like text. Moreover, we believe that reverse crossentropy, i.e., the cross-entropy of P relative to Q θ , is a better reflection of how a human would evaluate text generated by a model. Hence, we propose learning with MIXCE, an objective that mixes the forward and reverse crossentropies. We evaluate models trained with this objective on synthetic data settings (where P is known) and real data, and show that the resulting models yield better generated text without complex decoding strategies. https://github.com/bloomberg/ mixce-acl2023 * Work done during an internship at Bloomberg. 1 Unbiased sampling is vanilla random sampling, i.e., sampling with temperature=1.0. It is also called ancestral sampling (Eikema and Aziz, 2020) or pure sampling (Holtzman
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 被引用 9 次
- Language Generation with Strictly Proper Scoring RulesChenze Shao, Fandong Meng, Yijin Liu, Jie ZhouICML 2024 · 被引用 7 次
- Beyond MLE: Convex Learning for Text GenerationChenze Shao, Zhengrui Ma, Min Zhang, Yang FengNeurIPS 2023 · 被引用 5 次
- Compatibility-Aware Dynamic Fine-Tuning for Large Language ModelsYucheng Zhou, Junwei Sheng, Qianning Wang, Jianbing ShenACL 2026
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
相关 Paper
- Mixed Cross Entropy Loss for Neural Machine TranslationHaoran Li, Wei LuICML 2021 · 被引用 21 次
- SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with BacktrackingChris Cundy, Stefano ErmonICLR 2024 · 被引用 17 次
- Dual-objective Language Models: Training Efficiency Without OverfittingDavid Samuel, Lucas Georges Gabriel CharpentierICLR 2026
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 被引用 93 次
- Mixture of Inputs: Text Generation Beyond Discrete Token SamplingYufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang 等NeurIPS 2025
