Test Time Adaptation via Conjugate Pseudo-labels
Sachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico Kolter
Abstract
Test-time adaptation (TTA) refers to adapting neural networks to distribution shifts, with access to only the unlabeled test samples from the new domain at test-time. Prior TTA methods optimize over unsupervised objectives such as the entropy of model predictions in TENT [50], but it is unclear what exactly makes a good TTA loss. In this paper, we start by presenting a surprising phenomenon: if we attempt to meta-learn the "best" possible TTA loss over a wide class of functions, then we recover a function that is remarkably similar to (a temperature-scaled version of) the softmax-entropy employed by TENT. This only holds, however, if the classifier we are adapting is trained via cross-entropy loss; if the classifier is trained via squared loss, a different "best" TTA loss emerges. To explain this phenomenon, we analyze test-time adaptation through the lens of the training losses's convex conjugate. We show that under natural conditions, this (unsupervised) conjugate function can be viewed as a good local approximation to the original supervised loss and indeed, it recovers the "best" losses found by meta-learning. This leads to a generic recipe that can be used to find a good TTA loss for any given supervised training loss function of a general class. Empirically, our approach consistently dominates other TTA alternatives over a wide range of domain adaptation benchmarks. Our approach is particularly of interest when applied to classifiers trained with novel loss functions, e.g., the recently-proposed PolyLoss [25] function, where it differs substantially from (and outperforms) an entropy-based loss. Further, we show that our conjugate based approach can also be interpreted as a kind of self-training using a very specific soft label, which we refer to as the conjugate pseudo-label. Overall, our method provides a broad framework for better understanding and improving test-time adaptation. Code is available at https://github.com/locuslab/ tta_conjugate . Equal Contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60c40be1-28c8-4a1e-aa8a-a4352779c415Cited by top-tier papers47
- Test-Time Model Adaptation with Only Forward PassesShuaicheng Niu, Chunyan Miao, Guohao Chen, Pengcheng Wu et al.ICML 2024 · 77 citations
- On Pitfalls of Test-Time AdaptationHao Zhao, Yuejiang Liu, Alexandre Alahi, Tao LinICML 2023 · 72 citations
- RDumb: A simple approach that questions our progress in continual test-time adaptationOri Press, Steffen Schneider, Matthias Kümmerer, Matthias BethgeNeurIPS 2023 · 67 citations
- Towards Open-Set Test-Time Adaptation Utilizing the Wisdom of Crowds in Entropy MinimizationJungsoo Lee, Debasmit Das, Jaegul Choo, Sungha ChoiICCV 2023 · 48 citations
- On the Robustness of Open-World Test-Time Training: Self-Training with Dynamic Prototype ExpansionYushu Li, Xun Xu, Yongyi Su, Kui JiaICCV 2023 · 44 citations
Builds on30
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- Test-Time Adaptation via Self-Training with Nearest Neighbor InformationMinguk Jang, Sae-Young Chung, Hye Won ChungICLR 2023 · 12 citations
- Towards Understanding GD with Hard and Conjugate Pseudo-labels for Test-Time AdaptationJun-Kun Wang, Andre WibisonoICLR 2023 · 2 citations
- Towards Test Time Adaptation via Calibrated Entropy MinimizationHao Yang, Min Wang, Jinshen Jiang, Yun ZhouKDD 2024 · 3 citations
- Robust Mean Teacher for Continual and Gradual Test-Time AdaptationMario Döbler, Robert A. Marsden, Bin YangCVPR 2023
- CAFA: Class-Aware Feature Alignment for Test-Time AdaptationSanghun Jung, Jungsoo Lee, Nanhee Kim, Amirreza Shaban et al.ICCV 2023 · 23 citations
