Loss Functions and Operators Generated by f-Divergences
Vincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander, Mathieu Blondel
Abstract
The logistic loss (a.k.a. cross-entropy loss) is one of the most popular loss functions used for multiclass classification. It is also the loss function of choice for next-token prediction in language modeling. It is associated with the Kullback-Leibler (KL) divergence and the softargmax operator. In this work, we build upon Fenchel-Young losses to construct convex loss functions generated from f -divergences. Our loss functions generalize the logistic loss in two directions: i) by replacing the KL divergence with f -divergences and ii) by allowing non-uniform reference measures. We instantiate our framework for numerous f -divergences, recovering existing losses and creating new ones. By analogy with the logistic loss, the loss function generated by an f -divergence is associated with an operator, that we dub fsoftargmax. We derive a novel parallelizable bisection algorithm for computing the f -softargmax associated with any f -divergence. On the empirical side, one of the goals of this paper is to determine the effectiveness of loss functions beyond the classical cross-entropy in a language model setting, including on pre-training, post-training (SFT) and distillation. We show that the loss function generated by the α-divergence (which is equivalent to Tsallis α-negentropy in the case of unit reference measures) with α = 1.5 performs well across several tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8751d319-dafb-416a-8cc4-9912c043d90dCited by top-tier papers4
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token PredictionMathieu Blondel, Michael Sander, Germain Vivier-Ardisson, Tianlin Liu et al.ICML 2026 · 8 citations
- Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young LossesYuzhou Cao, Han Bao, Lei Feng, Bo AnNeurIPS 2025 · 4 citations
- Any-stepsize Gradient Descent for Separable Data under Fenchel-Young LossesHan Bao, Shinsaku Sakaue, Yuki TakezawaNeurIPS 2025 · 2 citations
- Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled RegularizationSafwan Labbi, Daniil Tiapkin, Paul Mangold, Eric MoulinesICLR 2026 · 2 citations
Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk et al.ICLR 2024 · 311 citations
Related papers
- Learning with Fitzpatrick LossesSeta Rakotomandimby, Jean-Philippe Chancelier, Michel De Lara, Mathieu BlondelNeurIPS 2024 · 6 citations
- -Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy DataJake Fawkes, Jason HartfordICML 2026
- Are All Losses Created Equal: A Neural Collapse PerspectiveJinxin Zhou, Chong You, Xiao Li, Kangning Liu et al.NeurIPS 2022 · 93 citations
- PolyLoss: A Polynomial Expansion Perspective of Classification Loss FunctionsZhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk et al.ICLR 2022 · 189 citations
- f-Divergence Based Classification: Beyond the Use of Cross-EntropyNicola Novello, Andrea M. TonelloICML 2024 · 18 citations
