Training for the Future: A Simple Gradient Interpolation Loss to Generalize Along Time
Anshul Nasery, Soumyadeep Thakur, Vihari Piratla, Abir De, Sunita Sarawagi
Abstract
In several real world applications, machine learning models are deployed to make predictions on data whose distribution changes gradually along time, leading to a drift between the train and test distributions. Such models are often re-trained on new data periodically, and they hence need to generalize to data not too far into the future. In this context, there is much prior work on enhancing temporal generalization, e.g. continuous transportation of past data, kernel smoothed time-sensitive parameters and more recently, adversarial learning of time-invariant features. However, these methods share several limitations, e.g, poor scalability, training instability, and dependence on unlabeled data from the future. Responding to the above limitations, we propose a simple method that starts with a model with time-sensitive parameters but regularizes its temporal complexity using a Gradient Interpolation (GI) loss. GI allows the decision boundary to change along time and can still prevent overfitting to the limited training time snapshots by allowing task-specific control over changes along time. We compare our method to existing baselines on multiple real-world datasets, which show that GI outperforms more complicated generative and adversarial approaches on the one hand, and simpler gradient regularization methods on the other.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6c62573-21ab-41d0-b3c7-0c444ee8f736Cited by top-tier papers21
- MADG: Margin-based Adversarial Learning for Domain GeneralizationAveen Dayal, Vimal K. B., Linga Reddy Cenkeramaddi, C. Krishna Mohan et al.NeurIPS 2023 · 102 citations
- Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular DataKai Helli, David Schnurr, Noah Hollmann, Samuel Müller et al.NeurIPS 2024 · 43 citations
- Learning Robust Spectral Dynamics for Temporal Domain GeneralizationEn Yu, Jie Lu, Xiaoyu Yang, Guangquan Zhang et al.NeurIPS 2025 · 22 citations
- Foresee What You Will Learn: Data Augmentation for Domain Generalization in Non-stationary EnvironmentQiuhao Zeng, Wei Wang, Fan Zhou, Charles Ling et al.AAAI 2023 · 22 citations
- Evolving Standardization for Continual Domain Generalization over Temporal DriftMixue Xie, Shuang Li, Longhui Yuan, Chi Harold Liu et al.NeurIPS 2023 · 21 citations
Builds on5
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 266 citations
- Efficient Domain Generalization via Common-Specific Low-Rank DecompositionVihari Piratla, Praneeth Netrapalli, Sunita SarawagiICML 2020 · 250 citations
- Continuously Indexed Domain AdaptationHao Wang, Hao He, Dina KatabiICML 2020 · 129 citations
- Learning to Adapt to Evolving DomainsHong Liu, Mingsheng Long, Jianmin Wang, Yu WangNeurIPS 2020 · 63 citations
- Follow the Perturbed Leader: Optimism and Fast Parallel Algorithms for Smooth Minimax GamesArun Sai Suggala, Praneeth NetrapalliNeurIPS 2020 · 22 citations
Related papers
- Temporal Generalization: A Reality CheckDivyam Madaan, Sumit Chopra, Kyunghyun ChoICLR 2026
- Temporal Domain Generalization with Drift-Aware Dynamic Neural NetworksGuangji Bai, Chen Ling, Liang ZhaoICLR 2023 · 6 citations
- CODA: Temporal Domain Generalization via Concept Drift SimulatorChia-Yuan Chang, Yu-Neng Chuang, Zhimeng Jiang, Kwei-Herng Lai et al.KDD 2025
- Latent Trajectory Learning for Limited Timestamps under Distribution Shift over TimeQiuhao Zeng, Changjian Shui, Long-Kai Huang, Peng Liu et al.ICLR 2024 · 15 citations
- Time-Varying Propensity Score to Bridge the Gap between the Past and PresentRasool Fakoor, Jonas Mueller, Zachary Chase Lipton, Pratik Chaudhari et al.ICLR 2024 · 4 citations
