Learning to Doubt: Forgetting Aware Learning for Neural Networks
Awanish Kumar, Soumyadeep Ghosh, Akshita Sharma, Rahul Gupta
Abstract
Modern neural networks are often miscalibrated, assigning high confidence to predictions they have not learned stably, which leads to overconfident errors under noise, imbalance, and distribution shift. We propose Forgetting-Aware Learning (FAL) using a simple ranking regularizer that leverages forgetting events during training as a temporal signal of epistemic uncertainty. The proposed regularizer penalizes high confidence on volatile samples, enforcing an isotonic relationship between confidence and learning stability and thereby reducing the mass of high-confidence mistakes that dominate calibration error. We show that FAL provides complementary benefits to Confidence-Aware Learning (CAL) using temporal instability as a regularization signal yielding a more self-aware model. Empirically, FAL improves calibration and confidence ranking while preserving accuracy. We also show that combining FAL and CAL with focal loss yields the strongest calibration in our study, significantly outperforming cross-entropy-based counterparts in the expected calibration error (ECE). We validate these findings across five popular image benchmark datasets, (with additional results on tabular data reported in the appendix), demonstrating that confidence should be earned by stable learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c142ccfe-d8fa-47be-8de6-10953c1f3488Related papers
- AdaFocal: Calibration-aware Adaptive Focal LossArindam Ghosh, Thomas Schaaf, Matthew GormleyNeurIPS 2022 · 71 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- Dual Focal Loss for CalibrationLinwei Tao, Minjing Dong, Chang XuICML 2023 · 56 citations
- Towards Understanding The Calibration Benefits of Sharpness-Aware MinimizationChengli Tan, Yubo Zhou, Haishan Ye, Guang Dai et al.ICLR 2026 · 3 citations
