Variance Reduced Training with Stratified Sampling for Forecasting Models
Yucheng Lu, Youngsuk Park, Lifan Chen, Yuyang Wang, Christopher De Sa, Dean P. Foster
Abstract
In large-scale time series forecasting, one often encounters the situation where the temporal patterns of time series, while drifting over time, differ from one another in the same dataset. In this paper, we provably show under such heterogeneity, training a forecasting model with commonly used stochastic optimizers (e.g. SGD) potentially suffers large variance on gradient estimation, and thus incurs long-time training. We show that this issue can be efficiently alleviated via stratification, which allows the optimizer to sample from pre-grouped time series strata. For better trading-off gradient variance and computation complexity, we further propose SCott (Stochastic Stratified Control Variate Gradient Descent), a variance reduced SGD-style optimizer that utilizes stratified sampling via control variate. In theory, we provide the convergence guarantee of SCott on smooth non-convex objectives. Empirically, we evaluate SCott and other baseline optimizers on both synthetic and real-world time series forecasting problems, and demonstrate SCott converges faster with respect to both iterations and wall clock time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 414b4130-6691-49c8-819a-0d3bb39b08feCited by top-tier papers4
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman et al.ICML 2024 · 45 citations
- GraB: Finding Provably Better Data Permutations than Random ReshufflingYucheng Lu, Wentao Guo, Christopher De SaNeurIPS 2022 · 23 citations
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao et al.ISSTA 2024 · 12 citations
- CD-GraB: Coordinating Distributed Example Orders for Provably Accelerated TrainingA. Feder Cooper, Wentao Guo, Khiem Pham, Tiancheng Yuan et al.NeurIPS 2023 · 9 citations
Builds on2
- Towards Fairness in Visual Recognition: Effective Strategies for Bias MitigationZeyu Wang, Klint Qinami, Ioannis Christos Karakozis, Kyle Genova et al.CVPR 2020
- Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing FlowsKashif Rasul, Abdul-Saboor Sheikh, Ingmar Schuster, Urs M. Bergmann et al.ICLR 2021
Related papers
- Parameter Decorrelation via Transition-Variance Alignment for Multivariate Time-series ForecastingJi-Eun Choi, Jae-Hong Lee, Joon Hyuk ChangICML 2026
- Spectral Subsampling MCMC for Stationary Time SeriesRobert Salomone, Matias Quiroz, Robert Kohn, Mattias Villani et al.ICML 2020 · 14 citations
- Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast ModelsEric Wang, Licheng Pan, Yuan Lu, Zi Ciu Chan et al.ICLR 2026 · 15 citations
- An Effective Hard Thresholding Method Based on Stochastic Variance Reduction for Nonconvex Sparse LearningGuannan Liang, Qianqian Tong, Chunjiang Zhu, Jinbo BiAAAI 2020 · 5 citations
- SCINet: Time Series Modeling and Forecasting with Sample Convolution and InteractionMinhao Liu, Ailing Zeng, Muxi Chen, Zhijian Xu et al.NeurIPS 2022 · 934 citations
