Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers
Qi Deng, Shuaicheng Niu, Ronghao Zhang, Yaofo Chen, Runhao Zeng, Jian Chen, Xiping Hu
Abstract
Test-time adaptation (TTA) aims to fine-tune a trained model online using unlabeled testing data to adapt to new environments or out-of-distribution data, demonstrating broad application potential in real-world scenarios. However, in this optimization process, unsupervised learning objectives like entropy minimization frequently encounter noisy learning signals. These signals produce unreliable gradients, which hinder the model's ability to converge to an optimal solution quickly and introduce significant instability into the optimization process. In this paper, we seek to resolve these issues from the perspective of optimizer design. Unlike prior TTA using manually designed optimizers like SGD, we employ a learning-to-optimize approach to automatically learn an optimizer, called Meta Gradient Generator (MGG). Specifically, we aim for MGG to effectively utilize historical gradient information during the online optimization process to optimize the current model. To this end, in MGG, we design a lightweight and efficient sequence modeling layer -gradient memory layer. It exploits a self-supervised reconstruction loss to compress historical gradient information into network parameters, thereby enabling better memorization ability over a long-term adaptation process. We only need a small number of unlabeled samples to pre-train MGG, and then the trained MGG can be deployed to process unseen samples. Promising results on ImageNet-C/R/Sketch/A indicate that our method surpasses current state-of-the-art methods with fewer updates, less data, and significantly shorter adaptation times. Compared with a previous SOTA SAR, we achieve 7.4% accuracy improvement and 4.2× faster adaptation speed on ImageNet-C. Code: https://github.com/keikeiqi/MGTTA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- PTTA: Purifying Malicious Samples for Test-Time Model AdaptationJing Ma, Hanlin Li, Xiang XiangICML 2025
- CONGA:Confidence-and-Gradient-Aware Learning Rate Schedule for Test Time AdaptationShaoran Lv, Xinyao Li, Jingjing LiICML 2026
- MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNormXiao Fan, Jingyan Jiang, Zhaoru Chen, Fanding Huang et al.AAAI 2026
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- MEMO: Test Time Robustness via Adaptation and AugmentationMarvin Zhang, Sergey Levine, Chelsea FinnNeurIPS 2022 · 595 citations
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
Related papers
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang et al.ICLR 2023
- Unified Entropy Optimization for Open-Set Test-Time AdaptationZhengqing Gao, Xu-Yao Zhang, Cheng-Lin LiuCVPR 2024
- Test Time Adaptation via Conjugate Pseudo-labelsSachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 152 citations
- MECTA: Memory-Economic Continual Test-Time Model AdaptationJunyuan Hong, Lingjuan Lyu, Jiayu Zhou, Michael SprangerICLR 2023
- Point-TTA: Test-Time Adaptation for Point Cloud Registration Using Multitask Meta-Auxiliary LearningAhmed Hatem, Yiming Qian, Yang WangICCV 2023 · 28 citations
