A Layer Selection Approach to Test Time Adaptation
Sabyasachi Sahoo, Mostafa ElAraby, Jonas Ngnawé, Yann Batiste Pequignot, Frédéric Precioso, Christian Gagné
Abstract
Test Time Adaptation (TTA) addresses the problem of distribution shift by adapting a pretrained model to a new domain during inference. When faced with challenging shifts, most methods collapse and perform worse than the original pretrained model. In this paper, we find that not all layers are equally receptive to the adaptation, and the layers with the most misaligned gradients often cause performance degradation. To address this, we propose GALA, a novel layer selection criterion to identify the most beneficial updates to perform during test time adaptation. This criterion can also filter out unreliable samples with noisy gradients. Its simplicity allows seamless integration with existing TTA loss functions, thereby preventing degradation and focusing adaptation on the most trainable layers. This approach also helps to regularize adaptation to preserve the pretrained features, which are crucial for handling unseen domains. Through extensive experiments, we demonstrate that the proposed layer selection framework improves the performance of existing TTA approaches across multiple datasets, domain shifts, model architectures, and TTA losses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c954905-8a99-46d3-b42e-7eb3dc06d3e7Cited by top-tier papers4
- ModHiFi: Identifying High Fidelity predictive components for Model ModificationDhruva Kashyap, Chaitanya Murti, Pranav K. Nayak, Tanay Narshana et al.NeurIPS 2025 · 1 citation
- Architecture-Agnostic Test-Time Adaptation via Backprop-Free Embedding AlignmentMA Xiao, Young D. Kwon, Pan Zhou, Dong MaICLR 2026
- A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to AdaptTomoya WakayamaICML 2026
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the EdgeYu Zhang, Xi Zhang, Hualin zhou, Xinyuan Chen et al.ICML 2026
Builds on51
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- CAFA: Class-Aware Feature Alignment for Test-Time AdaptationSanghun Jung, Jungsoo Lee, Nanhee Kim, Amirreza Shaban et al.ICCV 2023 · 23 citations
- PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time AdaptationSarthak Kumar Maharana, Baoming Zhang, Yunhui GuoAAAI 2025 · 7 citations
- What, How, and When Should Object Detectors Update in Continually Changing Test Domains?Jayeon Yoo, Dongkwan Lee, Inseop Chung, Donghyun Kim et al.CVPR 2024 · 10 citations
- Improved Test-Time Adaptation for Domain GeneralizationLiang Chen, Yong Zhang, Yibing Song, Ying Shan et al.CVPR 2023
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
