An Alternative Probabilistic Interpretation of the Huber Loss
Gregory P. Meyer
Abstract
The Huber loss is a robust loss function used for a wide range of regression tasks. To utilize the Huber loss, a parameter that controls the transitions from a quadratic function to an absolute value function needs to be selected. We believe the standard probabilistic interpretation that relates the Huber loss to the Huber density fails to provide adequate intuition for identifying the transition point. As a result, a hyper-parameter search is often necessary to determine an appropriate value. In this work, we propose an alternative probabilistic interpretation of the Huber loss, which relates minimizing the loss to minimizing an upperbound on the Kullback-Leibler divergence between Laplace distributions, where one distribution represents the noise in the ground-truth and the other represents the noise in the prediction. In addition, we show that the parameters of the Laplace distributions are directly related to the transition point of the Huber loss. We demonstrate, through a toy problem, that the optimal transition point of the Huber loss is closely related to the distribution of the noise in the ground-truth data. As a result, our interpretation provides an intuitive way to identify well-suited hyper-parameters by approximating the amount of noise in the data, which we demonstrate through a case study and experimentation on the Faster R-CNN and RetinaNet object detectors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ce29697-e62e-4c23-b922-0a3c1f4733b1Cited by top-tier papers12
- Iteratively Reweighted Least Squares for Basis Pursuit with Global Linear Convergence RateChristian Kümmerle, Claudio Mayrink Verdun, Dominik StögerNeurIPS 2021 · 25 citations
- Get Rid of Isolation: A Continuous Multi-task Spatio-Temporal Learning FrameworkZhongchao Yi, Zhengyang Zhou, Qihe Huang, Yanjiang Chen et al.NeurIPS 2024 · 18 citations
- PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal RepresentationYanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa, Simone Benatti et al.NeurIPS 2025 · 17 citations
- Variational Sparse Coding with Learned ThresholdingKion Fallah, Christopher J. RozellICML 2022 · 8 citations
- Beyond Structure: Invariant Crystal Property Prediction with Pseudo-Particle Ray DiffractionBin Cao, Yang Liu, Longhan Zhang, Yifan Wu et al.ICLR 2026 · 4 citations
Related papers
- Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler DivergenceXue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming et al.NeurIPS 2021 · 603 citations
- Probabilistic Regression for Visual TrackingMartin Danelljan, Luc Van Gool, Radu TimofteCVPR 2020
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 170 citations
- A Laplace-inspired Distribution on SO(3) for Probabilistic Rotation EstimationYingda Yin, Yang Wang, He Wang, Baoquan ChenICLR 2023 · 2 citations
- Effective Bayesian Heteroscedastic Regression with Deep Neural NetworksAlexander Immer, Emanuele Palumbo, Alexander Marx, Julia E. VogtNeurIPS 2023 · 34 citations
