Generalize for Future: Slow and Fast Trajectory Learning for CTR Prediction
Jian Zhu, Congcong Liu, Xue Jiang, Changping Peng, Zhangang Lin, Jingping Shao
Abstract
Deep neural networks (DNNs) have achieved significant advancements in click-through rate (CTR) prediction by demonstrating strong generalization on training data. However, in real-world scenarios, the assumption of independent and identically distributed (i.i.d.) conditions, which is fundamental to this problem, is often violated due to temporal distribution shifts. This violation can lead to suboptimal model performance when optimizing empirical risk without access to future data, resulting in overfitting on the training data and convergence to a single sharp minimum. To address this challenge, we propose a novel model updating framework called Slow and Fast Trajectory Learning (SFTL) network. SFTL aims to mitigate the discrepancy between past and future domains while quickly adapting to recent changes in small temporal drifts. This mechanism entails two interactions among three complementary learners: (i) the Working Learner, which updates model parameters using modern optimizers (e.g., Adam, Adagrad) and serves as the primary learner in the recommendation system, (ii) the Slow Learner, which is updated in each temporal domain by directly assigning the model weights of the working learner, and (iii) the Fast Learner, which is updated in each iteration by assigning exponentially moving average weights of the working learner. Additionally, we propose a novel rank-based trajectory loss to facilitate interaction between the working learner and trajectory learner, aiming to adapt to temporal drift and enhance performance in the current domain compared to the past. We provide theoretical understanding and conduct extensive experiments on real-world CTR prediction datasets to validate the effectiveness and efficiency of SFTL in terms of both convergence speed and model performance. The results demonstrate the superiority of SFTL over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Continuously Indexed Domain AdaptationHao Wang, Hao He, Dina KatabiICML 2020 · 129 citations
- How to Retrain Recommender System?: A Sequential Meta-Learning MethodYang Zhang, Fuli Feng, Chenxu Wang, Xiangnan He et al.SIGIR 2020 · 70 citations
Related papers
- Deep Time-Stream Framework for Click-through Rate Prediction by Tracking Interest EvolutionShu-Ting Shi, Wenhao Zheng, Jun Tang, Qing-Guo Chen et al.AAAI 2020 · 10 citations
- Reformulating CTR Prediction: Learning Invariant Feature Interactions for RecommendationYang Zhang, Tianhao Shi, Fuli Feng, Wenjie Wang et al.SIGIR 2023 · 20 citations
- Looking at CTR Prediction Again: Is Attention All You Need?Yuan Cheng, Yanbo XueSIGIR 2021 · 18 citations
- Improving Long-tail User CTR Prediction via Hierarchical Distribution AlignmentYifan Wang, Weizhi Ma, Min Zhang, Xiaoxiao Xu et al.KDD 2025
- Learning Fast and Slow for Online Time Series ForecastingQuang Pham, Chenghao Liu, Doyen Sahoo, Steven C. H. HoiICLR 2023 · 15 citations
