AETTA: Label-Free Accuracy Estimation for Test-Time Adaptation
Taeckyung Lee, Sorn Chottananurak, Taesik Gong, Sung-Ju Lee
Abstract
Test-time adaptation (TTA) has emerged as a viable solution to adapt pre-trained models to domain shifts using unlabeled test data. However, TTA faces challenges of adaptation failures due to its reliance on blind adaptation to unknown test samples in dynamic scenarios. Traditional methods for out-of-distribution performance estimation are limited by unrealistic assumptions in the TTA context, such as requiring labeled data or re-training models. To address this issue, we propose AETTA, a label-free accuracy estimation algorithm for TTA. We propose the prediction disagreement as the accuracy estimate, calculated by comparing the target model prediction with dropout inferences. We then improve the prediction disagreement to extend the applicability of AETTA under adaptation failures. Our extensive evaluation with four baselines and six TTA methods demonstrates that AETTA shows an average of 19.8%p more accurate estimation compared with the baselines. We further demonstrate the effectiveness of accuracy estimation with a model recovery case study, showcasing the practicality of our model recovery based on accuracy estimation. The source code is available at https://github.com/taeckyung/AETTA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9199ed3-9ce0-4827-b686-99991b0c0f6eCited by top-tier papers5
- Persistent Test-time Adaptation in Recurring Testing ScenariosTrung-Hieu Hoang, MinhDuc Vo, Minh DoNeurIPS 2024 · 20 citations
- Monitoring Risks in Test-Time AdaptationMona Schirmer, Metod Jazbec, Christian Andersson Naesseth, Eric T. NalisnickNeurIPS 2025 · 10 citations
- Automated Model Evaluation for Object Detection Via Prediction Consistency and ReliabilitySeungju Yoo, Hyuk Kwon, Joong-Won Hwang, Kibok LeeICCV 2025 · 1 citation
- AudioTest: Prioritizing Audio Test CasesYinghua Li, Xueqi Dang, Wendkûuni C. Ouédraogo, Jacques Klein et al.ISSTA 2025 · 1 citation
- FRET: Feature Redundancy Elimination for Test Time AdaptationLinjing You, Jiabao Lu, Xiayuan Huang, Xiangli NieICCV 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- MEMO: Test Time Robustness via Adaptation and AugmentationMarvin Zhang, Sergey Levine, Chelsea FinnNeurIPS 2022 · 595 citations
Related papers
- Label Shift Adapter for Test-Time Adaptation under Covariate and Label ShiftsSunghyun Park, Seunghan Yang, Jaegul Choo, Sungrack YunICCV 2023 · 28 citations
- PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time AdaptationSarthak Kumar Maharana, Baoming Zhang, Yunhui GuoAAAI 2025 · 7 citations
- Test-Time Adaptation with Binary FeedbackTaeckyung Lee, Sorn Chottananurak, Junsu Kim, Jinwoo Shin et al.ICML 2025
- CAFA: Class-Aware Feature Alignment for Test-Time AdaptationSanghun Jung, Jungsoo Lee, Nanhee Kim, Amirreza Shaban et al.ICCV 2023 · 23 citations
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
