Better Aggregation in Test-Time Augmentation
Divya Shanmugam, Davis W. Blalock, Guha Balakrishnan, John V. Guttag
Abstract
Test-time augmentation—the aggregation of predictions across transformed versions of a test input—is a common practice in image classification. Traditionally, predictions are combined using a simple average. In this paper, we present 1) experimental analyses that shed light on cases in which the simple average is suboptimal and 2) a method to address these shortcomings. A key finding is that even when test-time augmentation produces a net improvement in accuracy, it can change many correct predictions into incorrect predictions. We delve into when and why test-time augmentation changes a prediction from being correct to incorrect and vice versa. Building on these insights, we present a learning-based method for aggregating test-time augmentations. Experiments across a diverse set of models, datasets, and augmentations show that our method delivers consistent improvements over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e18f1d16-d67c-4afa-a88c-afbff9eda058Cited by top-tier papers31
- Diverse Data Augmentation with Diffusions for Effective Test-time Prompt TuningChun-Mei Feng, Kai Yu, Yong Liu, Salman Khan et al.ICCV 2023 · 172 citations
- POCO: Point Convolution for Surface ReconstructionAlexandre Boulch, Renaud MarletCVPR 2022 · 128 citations
- GraphPatcher: Mitigating Degree Bias for Graph Neural Networks via Test-time AugmentationMingxuan Ju, Tong Zhao, Wenhao Yu, Neil Shah et al.NeurIPS 2023 · 52 citations
- Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsMatteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini et al.NeurIPS 2024 · 47 citations
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim et al.AAAI 2023 · 34 citations
Builds on1
Related papers
- MEMO: Test Time Robustness via Adaptation and AugmentationMarvin Zhang, Sergey Levine, Chelsea FinnNeurIPS 2022 · 595 citations
- Test-time Augmentation Improves Efficiency in Conformal PredictionDivya Shanmugam, Helen Lu, Swami Sankaranarayanan, John V. GuttagCVPR 2025
- Combining Ensembles and Data Augmentation Can Harm Your CalibrationYeming Wen, Ghassen Jerfel, Rafael Muller, Michael W. Dusenberry et al.ICLR 2021 · 72 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Test-Time Training with Diversified Local Aggregation Consistency for Mortality Prediction using Clinical Time SeriesJingwen Xu, Fei Lyu, Pong C. YuenKDD 2025
