Learning with little mixing
Ingvar M. Ziemann, Stephen Tu
摘要
We study square loss in a realizable time-series framework with martingale difference noise. Our main result is a fast rate excess risk bound which shows that whenever a trajectory hypercontractivity condition holds, the risk of the least-squares estimator on dependent data matches the iid rate order-wise after a burn-in time. In comparison, many existing results in learning from dependent data have rates where the effective sample size is deflated by a factor of the mixing-time of the underlying process, even after the burn-in time. Furthermore, our results allow the covariate process to exhibit long range correlations which are substantially weaker than geometric ergodicity. We call this phenomenon learning with little mixing, and present several examples for when it occurs: bounded function classes for which the and norms are equivalent, ergodic finite state Markov chains, various parametric models, and a broad family of infinite dimensional ellipsoids. By instantiating our main result to system identification of nonlinear dynamics with generalized linear model transitions, we obtain a nearly minimax optimal excess risk bound after only a polynomial burn-in time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Transformers as Algorithms: Generalization and Stability in In-context LearningYingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, Samet OymakICML 2023 · 被引用 242 次
- From Self-Attention to Markov Models: Unveiling the Dynamics of Generative TransformersMuhammed Emrullah Ildiz, Yixiao Huang, Yingcong Li, Ankit Singh Rawat 等ICML 2024 · 被引用 45 次
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes 等NeurIPS 2023 · 被引用 42 次
- Streaming PCA for Markovian DataSyamantak Kumar, Purnamrita SarkarNeurIPS 2023 · 被引用 16 次
- Sharp Rates in Dependent Learning Theory: Avoiding Sample Size Deflation for the Square LossIngvar M. Ziemann, Stephen Tu, George J. Pappas, Nikolai MatniICML 2024 · 被引用 10 次
它引用的顶会 Paper4
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett 等ICML 2021 · 被引用 207 次
- Least Squares Regression with Markovian Data: Fundamental Limits and AlgorithmsDheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain 等NeurIPS 2020 · 被引用 73 次
- Near-optimal Offline and Streaming Algorithms for Learning Non-Linear Dynamical SystemsSuhas S. Kowshik, Dheeraj Nagaraj, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 被引用 28 次
- On Empirical Risk Minimization with Dependent and Heavy-Tailed DataAbhishek Roy, Krishnakumar Balasubramanian, Murat A. ErdogduNeurIPS 2021 · 被引用 22 次
相关 Paper
- Long-Context Linear System IdentificationOguz Kaan Yüksel, Mathieu Even, Nicolas FlammarionICLR 2025
- The noise level in linear regression with dependent dataIngvar M. Ziemann, Stephen Tu, George J. Pappas, Nikolai MatniNeurIPS 2023 · 被引用 7 次
- Prior Diffusiveness and Regret in the Linear-Gaussian BanditYifan Zhu, John Duchi, Benjamin Van RoyICML 2026 · 被引用 1 次
- Robust System Identification: Finite-sample Guarantees and Connection to RegularizationHyuk Park, Grani A. Hanasusanto, Yingying LiICLR 2025
- On the Consistency of Kernel Methods with Dependent ObservationsPierre-François Massiani, Sebastian Trimpe, Friedrich SolowjowICML 2024 · 被引用 2 次
