Loss Functions and Operators Generated by f-Divergences
Vincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander, Mathieu Blondel
摘要
The logistic loss (a.k.a. cross-entropy loss) is one of the most popular loss functions used for multiclass classification. It is also the loss function of choice for next-token prediction in language modeling. It is associated with the Kullback-Leibler (KL) divergence and the softargmax operator. In this work, we build upon Fenchel-Young losses to construct convex loss functions generated from f -divergences. Our loss functions generalize the logistic loss in two directions: i) by replacing the KL divergence with f -divergences and ii) by allowing non-uniform reference measures. We instantiate our framework for numerous f -divergences, recovering existing losses and creating new ones. By analogy with the logistic loss, the loss function generated by an f -divergence is associated with an operator, that we dub fsoftargmax. We derive a novel parallelizable bisection algorithm for computing the f -softargmax associated with any f -divergence. On the empirical side, one of the goals of this paper is to determine the effectiveness of loss functions beyond the classical cross-entropy in a language model setting, including on pre-training, post-training (SFT) and distillation. We show that the loss function generated by the α-divergence (which is equivalent to Tsallis α-negentropy in the case of unit reference measures) with α = 1.5 performs well across several tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token PredictionMathieu Blondel, Michael Sander, Germain Vivier-Ardisson, Tianlin Liu 等ICML 2026 · 被引用 8 次
- Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young LossesYuzhou Cao, Han Bao, Lei Feng, Bo AnNeurIPS 2025 · 被引用 4 次
- Any-stepsize Gradient Descent for Separable Data under Fenchel-Young LossesHan Bao, Shinsaku Sakaue, Yuki TakezawaNeurIPS 2025 · 被引用 2 次
- Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled RegularizationSafwan Labbi, Daniil Tiapkin, Paul Mangold, Eric MoulinesICLR 2026 · 被引用 2 次
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk 等ICLR 2024 · 被引用 311 次
相关 Paper
- Learning with Fitzpatrick LossesSeta Rakotomandimby, Jean-Philippe Chancelier, Michel De Lara, Mathieu BlondelNeurIPS 2024 · 被引用 6 次
- -Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy DataJake Fawkes, Jason HartfordICML 2026
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu 等NeurIPS 2022 · 被引用 93 次
- PolyLoss: A Polynomial Expansion Perspective of Classification Loss FunctionsZhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk 等ICLR 2022 · 被引用 189 次
- f-Divergence Based Classification: Beyond the Use of Cross-EntropyNicola Novello, Andrea M. TonelloICML 2024 · 被引用 18 次
