Instance-Conditional Timescales of Decay for Non-Stationary Learning
Nishant Jain, Pradeep Shenoy
Abstract
Slow concept drift is a ubiquitous, yet under-studied problem in practical machine learning systems. In such settings, although recent data is more indicative of future data, naively prioritizing recent instances runs the risk of losing valuable information from the past. We propose an optimization-driven approach towards balancing instance importance over large training windows. First, we model instance relevance using a mixture of multiple timescales of decay, allowing us to capture rich temporal trends. Second, we learn an auxiliary scorer model that recovers the appropriate mixture of timescales as a function of the instance itself. Finally, we propose a nested optimization objective for learning the scorer, by which it maximizes forward transfer for the learned model. Experiments on a large real-world dataset of 39M photos over a 9 year period show upto 15% relative gains in accuracy compared to other robust learning baselines. We replicate our gains on two collections of real-world datasets for non-stationary learning, and extend our work to continual learning settings where, too, we beat SOTA methods by large margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Learning model uncertainty as variance-minimizing instance weightsNishant Jain, Karthikeyan Shanmugam, Pradeep ShenoyICLR 2024 · 7 citations
- Improving Generalization via Meta-Learning on Hard SamplesNishant Jain, Arun Sai Suggala, Pradeep ShenoyCVPR 2024
Builds on10
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li et al.AAAI 2020 · 499 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Improving Out-of-Distribution Robustness via Selective AugmentationHuaxiu Yao, Yu Wang, Sai Li, Linjun Zhang et al.ICML 2022 · 275 citations
- Continual Prototype Evolution: Learning Online from Non-Stationary Data StreamsMatthias De Lange, Tinne TuytelaarsICCV 2021 · 251 citations
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
Related papers
- DeepBooTS: Dual-Stream Residual Boosting for Drift-Resilient Time-Series ForecastingDaojun Liang, Jing Chen, Xiao Wang, Yinglong Wang et al.AAAI 2026
- TRACE: A Generalizable Drift Detector for Streaming Data-Driven OptimizationYuan-Ting Zhong, Ting Huang, Xiaolin Xiao, Yue-Jiao GongAAAI 2026 · 1 citation
- Training for the Future: A Simple Gradient Interpolation Loss to Generalize Along TimeAnshul Nasery, Soumyadeep Thakur, Vihari Piratla, Abir De et al.NeurIPS 2021 · 42 citations
- Minimax Classification under Concept Drift with Multidimensional Adaptation and Performance GuaranteesVerónica Álvarez, Santiago Mazuelas, José Antonio LozanoICML 2022 · 6 citations
- The AdEMAMix Optimizer: Better, Faster, OlderMatteo Pagliardini, Pierre Ablin, David GrangierICLR 2025
