PatAug: Augmentation of Augmentation for Test-Time Adaptation
Xinyao Li, Dan Zhang, Zhekai Du, Lei Zhu, Zhi Chen, Jingjing Li
Abstract
The rich pretrained knowledge in vision-language models (VLMs) endows them with the ability to discriminate common objects given only category names, but may be challenged by out-of-distribution unlabeled samples. To address this limitation, test-time adaptation (TTA) dynamically adjusts VLMs to target distributions during inference. Current TTA frameworks rely heavily on unsupervised data augmentations to enhance sample informativeness, but remain vulnerable to naive augmented views. This work introduces Patch Augmentation (PatAug), a pixel-level perturbation framework that optimizes the benefits of informative augmentations and mitigates negative transformation impacts. Implemented as trainable pixels, PatAug are prepared given only category names before inference, introducing few additional overheads. The patches encode class-related semantic information. They assist VLMs in emphasizing on the compatible visual information in the images, restoring perturbed image details, while retaining unrecognized information. Such merits inspire the design of an augmentation of augmentation framework, where PatAug is applied to standard augmentation views for reliable TTA inference results. To better fit the target distributions, we adjust patches with a cross-modal similarity alignment loss and learnable patching weights. Experiments on natural and specialized domain shifts confirm the effectiveness of PatAug.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ea395a29-b85e-44d4-b957-6618be9b8c25Cited by top-tier papers2
- Generalizing Vision-Language Models with Dedicated Prompt GuidanceXinyao Li, Yinjie Min, Hongbo Chen, Zhekai Du et al.AAAI 2026
- USE: A Unified Self-Ensembling Framework for Test-Time Prompt TuningSiru Jiang, Jian Liang, Ran He, Tieniu TanICML 2026
Related papers
- RA-TTA: Retrieval-Augmented Test-Time Adaptation for Vision-Language ModelsYoungjun Lee, Doyoung Kim, Junhyeok Kang, Jihwan Bang et al.ICLR 2025
- Flatness Guided Test-Time Adaptation for Vision-Language ModelsAodi Li, Liansheng Zhuang, Xiao Long, Houqiang Li et al.ICLR 2026 · 1 citation
- Panda: Test-Time Adaptation with Negative Data AugmentationRuxi Deng, Wenxuan Bao, Tianxin Wei, Jingrui HeAAAI 2026 · 3 citations
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
- PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language ModelsSi Chen, Yujia Chen, Xiaotian Yin, Xin Liu et al.ACM MM 2025 · 1 citation
