Project and Probe: Sample-Efficient Adaptation by Interpolating Orthogonal Features
Annie S. Chen, Yoonho Lee, Amrith Setlur, Sergey Levine, Chelsea Finn
摘要
Transfer learning with a small amount of target data is an effective and common approach to adapting a pre-trained model to distribution shifts. In some situations, target data labels may be expensive to obtain, so we may only have access to a limited number of target data points. To make the most of a very small target dataset, we propose a lightweight, sample-efficient approach that learns a diverse set of features and adapts to a target distribution by interpolating these features. Our approach, PROJECT AND PROBE (PRO 2 ), first learns a linear projection that maps a pre-trained embedding onto orthogonal directions while being predictive of labels in the source dataset. The goal of this step is to learn a variety of predictive features, so that at least some of them remain useful after distribution shift. PRO 2 then learns a linear classifier on top of these projected features using a small target dataset. Theoretically, we find that PRO 2 results in more sample-efficient generalization by inducing a favorable bias-variance tradeoff. Our experiments on four datasets, with multiple distribution shift settings for each, show that PRO 2 improves performance by 5-15% when given limited target data compared to prior methods such as standard linear probing. we first use a source dataset to project pre-trained feature embeddings onto a smaller set of predictive features. We enforce pairwise orthogonality among all features, thereby ensuring that each projected dimension carries unique information not present in others. We expect this learned feature space to compactly contain a diverse set of predictive features while discarding information that is redundant or not predictive on the task. PRO 2 then uses the reduced set of features as a basis space for adaptation. Specifically, we fit a linear head on top of the projected embedding using labeled target data. Both the linear projection and the linear head require minimal computational overhead, making PRO 2 a practical method for adapting to new target distributions. Fig. 1 shows a visual summary of PRO 2 . To support our approach, we provide a theoretical analysis, in both a general setting with minimal distribution assumptions as well as the more specific setting of a shifted homoscedastic Gaussian model, showing how PRO 2 learns a projection matrix that results in better generalization due to a favorable bias-variance tradeoff. From this analysis, PRO 2 improves sample efficiency because it can learn useful, diverse features so that it is more likely to better recover the important directions for adaptation with a smaller projection dimension, allowing us to combat the variance introduced by a very small target dataset while maintaining low bias. We conduct experiments on a variety of distribution shift settings across 4 datasets. We find that standard linear probing, which is the default method used by prior works, is not the most data-efficient adaptation approach. Using PRO 2 , i.e. projecting with source data onto an informative feature-space basis and probing with target data, improves performance by 5-15% in few-shot adaptation to new target distributions. RELATED WORK Robustness and zero-shot generalization. Many prior works aim to improve robustness to various distribution shifts (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Investigating the Benefits of Projection Head for Representation LearningYihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 等ICLR 2024 · 被引用 23 次
- Optimizing importance weighting in the presence of sub-population shiftsFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICLR 2025
它引用的顶会 Paper27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
相关 Paper
- Few-shot Adaptation to Distribution Shifts By Mixing Source and Target EmbeddingsYihao Xue, Ali Payani, Yu Yang, Baharan MirzasoleimanICML 2024 · 被引用 4 次
- A Closer Look at the Few-Shot Adaptation of Large Vision-Language ModelsJulio Silva-Rodríguez, Sina Hajimiri, Ismail Ben Ayed, Jose DolzCVPR 2024 · 被引用 32 次
- Rethinking Few-Shot Adaptation of Vision-Language Models in Two StagesMatteo Farina, Massimiliano Mancini, Giovanni Iacca, Elisa RicciCVPR 2025
- Head2Toe: Utilizing Intermediate Representations for Better Transfer LearningUtku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. MozerICML 2022 · 被引用 103 次
- Structured Model Probing: Empowering Efficient Transfer Learning by Structured RegularizationZhi-Fan Wu, Chaojie Mao, Xue Wang, Jianwen Jiang 等CVPR 2024 · 被引用 1 次
