How Patterns Dictate Learnability in Sequential Data
Mario Morawski, Anaïs Després, Rémi Rehm
摘要
Sequential data - ranging from financial time series to natural language - has driven the growing adoption of autoregressive models. However, these algorithms rely on the presence of underlying patterns in the data, and their identification often depends heavily on human expertise. Misinterpreting these patterns can lead to model misspecification, resulting in increased generalization error and degraded performance. The recently proposed evolving pattern (EvoRate) metric addresses this by using the mutual information between the next data point and its past to guide regression order estimation and feature selection. Building on this idea, we introduce a general framework based on predictive information, defined as the mutual information between the past and the future, . This quantity naturally defines an information-theoretic learning curve, which quantifies the amount of predictive information available as the observation window grows. Using this formalism, we show that the presence or absence of temporal patterns fundamentally constrains the learnability of sequential models: even an optimal predictor cannot outperform the intrinsic information limit imposed by the data. We validate our framework through experiments on synthetic data, demonstrating its ability to assess model adequacy, quantify the inherent complexity of a dataset, and reveal interpretable structure in sequential data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series ForecastingKashif Rasul, Calvin Seward, Ingmar Schuster, Roland VollgrafICML 2021 · 被引用 500 次
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 被引用 243 次
- Predictive Information Accelerates Learning in RLKuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo 等NeurIPS 2020 · 被引用 82 次
- TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time SeriesArjun Ashok, Étienne Marcotte, Valentina Zantedeschi, Nicolas Chapados 等ICLR 2024 · 被引用 27 次
相关 Paper
- Towards Understanding Evolving Patterns in Sequential DataQiuhao Zeng, Long-Kai Huang, Qi Chen, Charles X. Ling 等NeurIPS 2024 · 被引用 5 次
- Enhancing Evolving Domain Generalization through Dynamic Latent RepresentationsBinghui Xie, Yongqiang Chen, Jiaqi Wang, Kaiwen Zhou 等AAAI 2024 · 被引用 9 次
- Optimal Information Retention for Time-Series ExplanationsJinghang Yue, Jing Wang, Lu Zhang, Shuo Zhang 等ICML 2025
- Learning Time-Aware Causal Representation for Model Generalization in Evolving DomainsZhuo He, Shuang Li, Wenze Song, Longhui Yuan 等ICML 2025
- Quantifying Semantic Emergence in Language ModelsHang Chen, Xinyu Yang, Jiaying Zhu, Wenya WangACL 2025
