Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
Liam Collins, Hamed Hassani, Mahdi Soltanolkotabi, Aryan Mokhtari, Sanjay Shakkottai
Abstract
An increasingly popular machine learning paradigm is to pretrain a neural network (NN) on many tasks offline, then adapt it to downstream tasks, often by re-training only the last linear layer of the network. This approach yields strong downstream performance in a variety of contexts, demonstrating that multitask pretraining leads to effective feature learning. Although several recent theoretical studies have shown that shallow NNs learn meaningful features when either (i) they are trained on a single task or (ii) they are linear, very little is known about the closer-to-practice case of nonlinear NNs trained on multiple tasks. In this work, we present the first results proving that feature learning occurs during training with a nonlinear model on multiple tasks. Our key insight is that multi-task pretraining induces a pseudo-contrastive loss that favors representations that align points that typically have the same label across tasks. Using this observation, we show that when the tasks are binary classification tasks with labels depending on the projection of the data onto an r -dimensional subspace within the d ≫ r -dimensional input space, a simple gradient-based multitask learning algorithm on a two-layer ReLU NN recovers this projection, allowing for generalization to downstream tasks with sample and neuron complexity independent of d . In contrast, we show that with high probability over the draw of a single task, training on this single task cannot guarantee to learn all r ground-truth features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural NetworksBehrad Moniri, Donghwan Lee, Hamed Hassani, Edgar DobribanICML 2024 · 38 citations
- Sample-Efficient Linear Representation Learning from Non-IID Non-Isotropic DataThomas T. C. K. Zhang, Leonardo Felipe Toso, James Anderson, Nikolai MatniICLR 2024 · 17 citations
- Neural Collapse in Multi-Task LearningYoujun Wang, Boqi Li, Xin Zou, Weiwei LiuICLR 2026 · 16 citations
- Uncertainty-aware Graph-based Hyperspectral Image ClassificationLinlin Yu, Yifei Lou, Feng ChenICLR 2024 · 9 citations
- Adversarially Robust Multi-task Representation LearningAustin Watkins, Thanh Nguyen-Tang, Enayat Ullah, Raman AroraNeurIPS 2024 · 5 citations
Builds on26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 1,081 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 218 citations
Related papers
- Inductive biases of multi-task learning and finetuning: multiple regimes of feature reuseSamuel Lippl, Jack W. LindseyNeurIPS 2024 · 13 citations
- Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU NetworksDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2025
- The Trade-off between Universality and Label Efficiency of Representations from Contrastive LearningZhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram et al.ICLR 2023 · 1 citation
- Features are fate: a theory of transfer learning in high-dimensional regressionJavan Tahir, Surya Ganguli, Grant M. RotskoffICML 2025
- Transfer Learning in Infinite Width Feature Learning NetworksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICLR 2026 · 2 citations
