Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal Features
Annie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine, Chelsea Finn
Abstract
Transfer learning with a small amount of target data is an effective and common approach to adapting a pre-trained model to distribution shifts. In some situations, target data labels may be expensive to obtain, so we may only have access to a limited number of target data points. To make the most of a very small target dataset, we propose a lightweight, sample-efficient approach that learns a diverse set of features and adapts to a target distribution by interpolating these features. Our approach, PROJECT AND PROBE (PRO 2 ), first learns a linear projection that maps a pre-trained embedding onto orthogonal directions while being predictive of labels in the source dataset. The goal of this step is to learn a variety of predictive features, so that at least some of them remain useful after distribution shift. PRO 2 then learns a linear classifier on top of these projected features using a small target dataset. Theoretically, we find that PRO 2 results in more sample-efficient generalization by inducing a favorable bias-variance tradeoff. Our experiments on four datasets, with multiple distribution shift settings for each, show that PRO 2 improves performance by 5-15% when given limited target data compared to prior methods such as standard linear probing. we first use a source dataset to project pre-trained feature embeddings onto a smaller set of predictive features. We enforce pairwise orthogonality among all features, thereby ensuring that each projected dimension carries unique information not present in others. We expect this learned feature space to compactly contain a diverse set of predictive features while discarding information that is redundant or not predictive on the task. PRO 2 then uses the reduced set of features as a basis space for adaptation. Specifically, we fit a linear head on top of the projected embedding using labeled target data. Both the linear projection and the linear head require minimal computational overhead, making PRO 2 a practical method for adapting to new target distributions. Fig. 1 shows a visual summary of PRO 2 . To support our approach, we provide a theoretical analysis, in both a general setting with minimal distribution assumptions as well as the more specific setting of a shifted homoscedastic Gaussian model, showing how PRO 2 learns a projection matrix that results in better generalization due to a favorable bias-variance tradeoff. From this analysis, PRO 2 improves sample efficiency because it can learn useful, diverse features so that it is more likely to better recover the important directions for adaptation with a smaller projection dimension, allowing us to combat the variance introduced by a very small target dataset while maintaining low bias. We conduct experiments on a variety of distribution shift settings across 4 datasets. We find that standard linear probing, which is the default method used by prior works, is not the most data-efficient adaptation approach. Using PRO 2 , i.e. projecting with source data onto an informative feature-space basis and probing with target data, improves performance by 5-15% in few-shot adaptation to new target distributions. RELATED WORK Robustness and zero-shot generalization. Many prior works aim to improve robustness to various distribution shifts (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5313bd67-b84b-47ba-852c-22a11002eba7Cited by top-tier papers2
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi et al.ICLR 2024 · 23 citations
- Optimizing importance weighting in the presence of sub-population shiftsFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICLR 2025
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
Related papers
- Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsYihao Xue, Ali Payani, Yu Yang, Baharan MirzasoleimanICML 2024 · 4 citations
- A Closer Look at the Few-Shot Adaptation of Large Vision-Language ModelsJulio Silva-Rodríguez, Sina Hajimiri, Ismail Ben Ayed, Jose DolzCVPR 2024 · 32 citations
- Rethinking Few-Shot Adaptation of Vision-Language Models in Two StagesMatteo Farina, Massimiliano Mancini, Giovanni Iacca, Elisa RicciCVPR 2025
- Head2Toe: Utilizing Intermediate Representations for Better Transfer LearningUtku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. MozerICML 2022 · 103 citations
- Structured Model Probing: Empowering Efficient Transfer Learning by Structured RegularizationZhi-Fan Wu, Chaojie Mao, Xue Wang, Jianwen Jiang et al.CVPR 2024 · 1 citation
