On the benefits of maximum likelihood estimation for Regression and Forecasting
Pranjal Awasthi, Abhimanyu Das, Rajat Sen, Ananda Theertha Suresh
Abstract
We advocate for a practical Maximum Likelihood Estimation (MLE) approach towards designing loss functions for regression and forecasting, as an alternative to the typical approach of direct empirical risk minimization on a specific target metric. The MLE approach is better suited to capture inductive biases such as prior domain knowledge in datasets, and can output post-hoc estimators at inference time that can optimize different types of target metrics. We present theoretical results to demonstrate that our approach is competitive with any estimator for the target metric under some general conditions. In two example practical settings, Poisson and Pareto regression, we show that our competitive results can be used to prove that the MLE approach has better excess risk bounds than directly minimizing the target metric. We also demonstrate empirically that our method instantiated with a well-designed general purpose mixture likelihood family can obtain superior performance for a variety of tasks across time-series forecasting and regression datasets with different data distributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5896522c-ee4b-4f1d-b043-306fc0825e5cCited by top-tier papers4
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
- Unified Training of Universal Time Series Forecasting TransformersGerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong et al.ICML 2024 · 513 citations
- Neural Spline Search for Quantile Probabilistic ModelingRuoxi Sun, Chun-Liang Li, Sercan Ö. Arik, Michael W. Dusenberry et al.AAAI 2023 · 5 citations
- Estimating Unknown Population Sizes Using the Hypergeometric DistributionLiam Hodgson, Danilo BzdokICML 2024
Builds on3
- Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural NetworksZonghan Wu, Shirui Pan, Guodong Long, Jing Jiang et al.KDD 2020 · 1,738 citations
- N-BEATS: Neural basis expansion analysis for interpretable time series forecastingBoris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua BengioICLR 2020 · 1,550 citations
- Efficient First-Order Contextual Bandits: Prediction, Allocation, and Triangular DiscriminationDylan J. Foster, Akshay KrishnamurthyNeurIPS 2021 · 62 citations
Related papers
- Empirical Gaussian ProcessesJihao Andreas Lin, Sebastian Ament, Louis Tiao, David Eriksson et al.ICML 2026
- Calibration by Distribution Matching: Trainable Kernel Calibration MetricsCharlie Marx, Sofian Zalouk, Stefano ErmonNeurIPS 2023 · 21 citations
- Quantile Risk Control: A Flexible Framework for Bounding the Probability of High-Loss PredictionsJake Snell, Thomas P. Zollo, Zhun Deng, Toniann Pitassi et al.ICLR 2023 · 1 citation
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang et al.ICCV 2021 · 286 citations
- Maximum Likelihood Estimation is All You Need for Well-Specified Covariate ShiftJiawei Ge, Shange Tang, Jianqing Fan, Cong Ma et al.ICLR 2024 · 16 citations
