Towards Understanding GD with Hard and Conjugate Pseudo-labels for Test-Time Adaptation
Jun-Kun Wang, Andre Wibisono
摘要
We consider a setting that a model needs to adapt to a new domain under distribution shifts, given that only unlabeled test samples from the new domain are accessible at test time. A common idea in most of the related works is constructing pseudo-labels for the unlabeled test samples and applying gradient descent (GD) to a loss function with the pseudo-labels. Recently, propose conjugate labels, which is a new kind of pseudo-labels for self-training at test time. They empirically show that the conjugate label outperforms other ways of pseudo-labeling on many domain adaptation benchmarks. However, provably showing that GD with conjugate labels learns a good classifier for test-time adaptation remains open. In this work, we aim at theoretically understanding GD with hard and conjugate labels for a binary classification problem. We show that for square loss, GD with conjugate labels converges to an -optimal predictor under a Gaussian model for any arbitrarily small , while GD with hard pseudo-labels fails in this task. We also analyze them under different loss functions for the update. Our results shed lights on understanding when and why GD with hard labels or conjugate labels works in test-time adaptation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Protected Test-Time Adaptation via Online Entropy Matching: A Betting ApproachYarin Bar, Shalev Shaer, Yaniv RomanoNeurIPS 2024 · 被引用 27 次
- Multi-Modal Continual Test-Time Adaptation for 3D Semantic SegmentationHaozhi Cao, Yuecong Xu, Jianfei Yang, Pengyu Yin 等ICCV 2023 · 被引用 27 次
- Uncovering Adversarial Risks of Test-Time AdaptationTong Wu, Feiran Jia, Xiangyu Qi, Jiachen T. Wang 等ICML 2023 · 被引用 12 次
- Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-LineEungyeup Kim, Mingjie Sun, Christina Baek, Aditi Raghunathan 等NeurIPS 2024 · 被引用 12 次
- Towards Understanding Extrapolation: a Causal LensLingjing Kong, Guangyi Chen, Petar Stojanov, Haoxuan Li 等NeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper26
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Test Time Adaptation via Conjugate Pseudo-labelsSachin Goyal, Mingjie Sun, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 被引用 152 次
- Contrastive Test-Time AdaptationDian Chen, Dequan Wang, Trevor Darrell, Sayna EbrahimiCVPR 2022 · 被引用 219 次
- STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language ModelsYiwei Fu, Hui Wan, Xiao Luo, Minghua DengCVPR 2026
- TTA-FedDG: Leveraging Test-Time Adaptation to Address Federated Domain GeneralizationHaoyuan Liang, Xinyu Zhang, Shilei Cao, Guowen Li 等AAAI 2025 · 被引用 4 次
- PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time AdaptationSarthak Kumar Maharana, Baoming Zhang, Yunhui GuoAAAI 2025 · 被引用 7 次
