Emo: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
Siyu Ren, Zhiyong Wu, Kenny Q. Zhu
Abstract
Neural language models are probabilistic models of human text. They are predominantly trained using maximum likelihood estimation (MLE), which is equivalent to minimizing the forward cross-entropy between the empirical data distribution and the model distribution. However, various degeneration phenomena are still widely observed when decoding from the distributions learned by such models. We establish that the forward cross-entropy is suboptimal as a distance metric for aligning human and model distribution due to its (1) recall-prioritization (2) negative diversity ignorance and (3) train-test mismatch. In this paper, we propose Earth Mover Distance Optimization (EMO) for auto-regressive language modeling. EMO capitalizes on the inherent properties of earth mover distance to address the aforementioned challenges. Due to the high complexity of direct computation, we further introduce a feasible upper bound for EMO to ease end-to-end training. Upon extensive evaluation of language models trained using EMO and MLE. We find that EMO demonstrates a consistently better language modeling performance than MLE across domains. Moreover, EMO demonstrates noteworthy enhancements in downstream performance with minimal fine-tuning on merely 25,000 sentences. This highlights the tremendous potential of EMO as a lightweight calibration method for enhancing large-scale pre-trained language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bf7992d-edd2-46a0-bca2-900985501e68Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
Related papers
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu et al.EMNLP 2020 · 11 citations
- Tailoring Language Generation Models under Total Variation DistanceHaozhe Ji, Pei Ke, Zhipeng Hu, Rongsheng Zhang et al.ICLR 2023 · 2 citations
- Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMsYuxiao Lu, Arunesh Sinha, Pradeep VarakanthamICLR 2025
- EMLight: Lighting Estimation via Spherical Distribution ApproximationFangneng Zhan, Changgong Zhang, Yingchen Yu, Yuan Chang et al.AAAI 2021 · 73 citations
- Calibration, Entropy Rates, and Memory in Language ModelsMark Braverman, Xinyi Chen, Sham M. Kakade, Karthik Narasimhan et al.ICML 2020 · 48 citations
