MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
Shiyue Zhang, Shijie Wu, Ozan Irsoy, Steven Lu, Mohit Bansal, Mark Dredze, David S. Rosenberg
Abstract
Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q θ relative to the data distribution Pthat is, minimizing the forward cross-entropy, which is equivalent to maximum likelihood estimation (MLE). We have observed that models trained in this way may "over-generalize", in the sense that they produce non-human-like text. Moreover, we believe that reverse crossentropy, i.e., the cross-entropy of P relative to Q θ , is a better reflection of how a human would evaluate text generated by a model. Hence, we propose learning with MIXCE, an objective that mixes the forward and reverse crossentropies. We evaluate models trained with this objective on synthetic data settings (where P is known) and real data, and show that the resulting models yield better generated text without complex decoding strategies. https://github.com/bloomberg/ mixce-acl2023 * Work done during an internship at Bloomberg. 1 Unbiased sampling is vanilla random sampling, i.e., sampling with temperature=1.0. It is also called ancestral sampling (Eikema and Aziz, 2020) or pure sampling (Holtzman
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 779aa48d-4cc2-4649-85d0-a0118c22a0c6Cited by top-tier papers4
- Emo: Earth Mover Distance Optimization for Auto-Regressive Language ModelingSiyu Ren, Zhiyong Wu, Kenny Q. ZhuICLR 2024 · 9 citations
- Language Generation with Strictly Proper Scoring RulesChenze Shao, Fandong Meng, Yijin Liu, Jie ZhouICML 2024 · 7 citations
- Beyond MLE: Convex Learning for Text GenerationChenze Shao, Zhengrui Ma, Min Zhang, Yang FengNeurIPS 2023 · 5 citations
- Compatibility-Aware Dynamic Fine-Tuning for Large Language ModelsYucheng Zhou, Junwei Sheng, Qianning Wang, Jianbing ShenACL 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
Related papers
- Mixed Cross Entropy Loss for Neural Machine TranslationHaoran Li, Wei LuICML 2021 · 21 citations
- SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with BacktrackingChris Cundy, Stefano ErmonICLR 2024 · 17 citations
- Dual-objective Language Models: Training Efficiency Without OverfittingDavid Samuel, Lucas Georges Gabriel CharpentierICLR 2026
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
- Mixture of Inputs: Text Generation Beyond Discrete Token SamplingYufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang et al.NeurIPS 2025
