Sample-Efficient Linear Representation Learning from Non-IID Non-Isotropic Data
Thomas T. C. K. Zhang, Leonardo Felipe Toso, James Anderson, Nikolai Matni
Abstract
A powerful concept behind much of the recent progress in machine learning is the extraction of common features across data from heterogeneous sources or tasks. Intuitively, using all of one's data to learn a common representation function benefits both computational effort and statistical generalization by leaving a smaller number of parameters to fine-tune on a given task. Toward theoretically grounding these merits, we propose a general setting of recovering linear operators from noisy vector measurements , where the covariates may be both non-i.i.d. and non-isotropic. We demonstrate that existing isotropy-agnostic representation learning approaches incur biases on the representation update, which causes the scaling of the noise terms to lose favorable dependence on the number of source tasks. This in turn can cause the sample complexity of representation learning to be bottlenecked by the single-task data size. We introduce an adaptation, \texttt{De-bias&Feature-Whiten} (), of the popular alternating minimization-descent scheme proposed independently in Collins et al., (2021) and Nayer and Vaswani (2022), and establish linear convergence to the optimal representation with noise level scaling down with the source data size. This leads to generalization bounds on the same order as an oracle empirical risk minimizer. We verify the vital importance of on various numerical simulations. In particular, we show that vanilla alternating-minimization descent fails catastrophically even for iid, but mildly non-isotropic data. Our analysis unifies and generalizes prior work, and provides a flexible framework for a wider range of applications, such as in controls and dynamical systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Regret Analysis of Multi-task Representation Learning for Linear-Quadratic Adaptive ControlBruce D. Lee, Leonardo F. Toso, Thomas T. C. K. Zhang, James Anderson et al.AAAI 2025 · 4 citations
- Guarantees for Nonlinear Representation Learning: Non-identical Covariates, Dependent Data, Fewer SamplesThomas T. C. K. Zhang, Bruce D. Lee, Ingvar M. Ziemann, George J. Pappas et al.ICML 2024 · 2 citations
- On the Power of Source Screening for Learning Shared Feature ExtractorsMuxing Wang, Connor Mclaughlin, Lili SuICML 2026
- On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature LearningThomas T. C. K. Zhang, Behrad Moniri, Ansh Nagwekar, Faraz Rahman et al.ICML 2025
Builds on8
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 1,081 citations
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 218 citations
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 209 citations
- Meta-Adaptive Nonlinear Control: Theory and AlgorithmsGuanya Shi, Kamyar Azizzadenesheli, Michael O'Connell, Soon-Jo Chung et al.NeurIPS 2021 · 62 citations
- Provable Benefit of Multitask Representation Learning in Reinforcement LearningYuan Cheng, Songtao Feng, Jing Yang, Hong Zhang et al.NeurIPS 2022 · 33 citations
Related papers
- Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual BanditsJiabin Lin, Shana Moothedath, Namrata VaswaniICML 2024 · 9 citations
- A Distribution-dependent Analysis of Meta LearningMikhail Konobeev, Ilja Kuzborskij, Csaba SzepesváriICML 2021 · 6 citations
- Gradient Extrapolation for Debiased Representation LearningIhab Asaad, Maha Shadaydeh, Joachim DenzlerICCV 2025 · 4 citations
- Statistically and Computationally Efficient Linear Meta-representation LearningKiran Koshy Thekumparampil, Prateek Jain, Praneeth Netrapalli, Sewoong OhNeurIPS 2021 · 28 citations
- Efficient Alternating Minimization with Applications to Weighted Low Rank ApproximationZhao Song, Mingquan Ye, Junze Yin, Lichen ZhangICLR 2025 · 1 citation
