Test Time Adaptation via Conjugate Pseudo-labels
Sachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico Kolter
摘要
Test-time adaptation (TTA) refers to adapting neural networks to distribution shifts, with access to only the unlabeled test samples from the new domain at test-time. Prior TTA methods optimize over unsupervised objectives such as the entropy of model predictions in TENT [50], but it is unclear what exactly makes a good TTA loss. In this paper, we start by presenting a surprising phenomenon: if we attempt to meta-learn the "best" possible TTA loss over a wide class of functions, then we recover a function that is remarkably similar to (a temperature-scaled version of) the softmax-entropy employed by TENT. This only holds, however, if the classifier we are adapting is trained via cross-entropy loss; if the classifier is trained via squared loss, a different "best" TTA loss emerges. To explain this phenomenon, we analyze test-time adaptation through the lens of the training losses's convex conjugate. We show that under natural conditions, this (unsupervised) conjugate function can be viewed as a good local approximation to the original supervised loss and indeed, it recovers the "best" losses found by meta-learning. This leads to a generic recipe that can be used to find a good TTA loss for any given supervised training loss function of a general class. Empirically, our approach consistently dominates other TTA alternatives over a wide range of domain adaptation benchmarks. Our approach is particularly of interest when applied to classifiers trained with novel loss functions, e.g., the recently-proposed PolyLoss [25] function, where it differs substantially from (and outperforms) an entropy-based loss. Further, we show that our conjugate based approach can also be interpreted as a kind of self-training using a very specific soft label, which we refer to as the conjugate pseudo-label. Overall, our method provides a broad framework for better understanding and improving test-time adaptation. Code is available at https://github.com/locuslab/ tta_conjugate . Equal Contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- Test-Time Model Adaptation with Only Forward PassesShuaicheng Niu, Chunyan Miao, Guohao Chen, Pengcheng Wu 等ICML 2024 · 被引用 77 次
- On Pitfalls of Test-Time AdaptationHao Zhao, Yuejiang Liu, Alexandre Alahi, Tao LinICML 2023 · 被引用 72 次
- RDumb: A simple approach that questions our progress in continual test-time adaptationOri Press, Steffen Schneider, Matthias Kümmerer, Matthias BethgeNeurIPS 2023 · 被引用 67 次
- Towards Open-Set Test-Time Adaptation Utilizing the Wisdom of Crowds in Entropy MinimizationJungsoo Lee, Debasmit Das, Jaegul Choo, Sungha ChoiICCV 2023 · 被引用 48 次
- On the Robustness of Open-World Test-Time Training: Self-Training with Dynamic Prototype ExpansionYushu Li, Xun Xu, Yongyi Su, Kui JiaICCV 2023 · 被引用 44 次
它引用的顶会 Paper30
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
相关 Paper
- Test-Time Adaptation via Self-Training with Nearest Neighbor InformationMinguk Jang, Sae-Young Chung, Hye Won ChungICLR 2023 · 被引用 12 次
- Towards Understanding GD with Hard and Conjugate Pseudo-labels for Test-Time AdaptationJun-Kun Wang, Andre WibisonoICLR 2023 · 被引用 2 次
- Towards Test Time Adaptation via Calibrated Entropy MinimizationHao Yang, Min Wang, Jinshen Jiang, Yun ZhouKDD 2024 · 被引用 3 次
- Robust Mean Teacher for Continual and Gradual Test-Time AdaptationMario Döbler, Robert A. Marsden, Bin YangCVPR 2023
- CAFA: Class-Aware Feature Alignment for Test-Time AdaptationSanghun Jung, Jungsoo Lee, Nanhee Kim, Amirreza Shaban 等ICCV 2023 · 被引用 23 次
